BenchmarksIndustrial · Industrial 162
IndustrialIndustry benchmark

Industrial-162

Updated: 7/22/2026

Tests whether an agent can formulate and solve industrial-scale combinatorial optimization problems - returning verified optimal solutions, or proofs of infeasibility, with solver tools.

162 problems·4 systems

Pramaana scores 100% across everything

All 162 problems solved - every optimum verified, every infeasible instance correctly proven. Pramaana's formalized domains leave nothing on the table.

Problems
162
Domains
22
INDUSTRIAL AGENT VERIFIED · PRAMAANA LABS100%VERIFIED

Results

#SystemProblems solvedSolved
1
PramaanaPramaana Labs
100.0% ± 0.0162/162
2
Claude Fable 5 (max)AnthropicBest frontier model
4.9% ± 1.78/162
3
GPT-5.6 Sol (xhigh)OpenAI
2.5% ± 1.24/162
4
Kimi K3Moonshot AI
2.5% ± 1.24/162

A problem counts as solved only when the system's formulation runs within the 600-second solver budget and returns a valid optimal solution - or correctly proves infeasibility. Refusals count as failures - Claude Fable 5 declined 7 biology-related problems.

Composition by domain

22 domains

Logistics
38
Combinatorial
16
Manufacturing
13
Telecom
11
Bio
10
Computing
10
Energy
10
Workforce
9
Finance
6
Packing
6
Airline
5
Defense
5
Media
5
Rail
5
Chemical
3
Security
3
Education
2
Construction
1
Forestry
1
ML
1
Network
1
Procurement
1

162 mixed-integer optimization problems across 22 industrial domains, drawn from the MIPLIB 2017 reference library.

Overview

Scheduling industrial and logistics operations is a complex combinatorial optimization problem. A single instance can hold hundreds of thousands of variables that must be tuned to compute the optimal schedule, and even minor gains compound across multi-trillion dollar logistics networks into billions in savings. Yet logistics and scheduling remain in the grips of Good Old-Fashioned AI.

Industrial-162 collects 162 mixed-integer optimization problems spanning 22 domains, drawn from the MIPLIB 2017 reference library. Each system gets solver tools and a 600-second budget per problem; a problem counts as solved only when the formulation returns a valid optimal solution - or correctly proves the instance infeasible.

We formalized the logistics domains. Pramaana solves the full portfolio while frontier LLMs - even when equipped with solver tools - struggle to generate optimal valid solutions.

Open the MIPLIB 2017 instance library

Key takeaways

  • Frontier LLMs struggle to code optimization domains. Claude Fable 5, GPT-5.6 Sol, and Kimi K3 each solve fewer than 5% of the portfolio. Most fail to generate the right formulation code, and much of the generated code is bloated enough to blow the solver's 600-second timeout.
  • Frontier capability is narrow and convergent. Kimi K3 and GPT-5.6 Sol pass the identical 4 problems out of 162; Claude Fable 5 passes a strict superset with 4 more. Capability looks limited to the small subset of problems encountered in training.
  • Model sovereignty cuts into coverage. Claude Fable 5 - the leading LLM here - refused 7 biology-related problems as a biosafety risk. The filters appear too broad, missing the substance of the problems.

Benchmark scope

  • Formulate the scheduling or allocation problem as a mixed-integer program
  • Generate solver code and run it within a 600-second execution budget
  • Return a valid optimal solution - or correctly prove the instance infeasible
  • Cover 22 domains - logistics, manufacturing, telecom, energy, workforce, and more

Updates

Industrial-162

Launch portfolio: 162 optimization problems across 22 industrial domains, evaluated on three frontier models with solver tools against Pramaana's formalized domain baseline. Pramaana solves all 162.