BenchmarksNewsAbout
Benchmark catalog

Real-world benchmarks for long-horizon agent work.

Three benchmarks across legal, healthcare, and industrial optimization, each measured against Pramaana's formalized baseline. Pramaana scores 100% on every benchmark - the best frontier models pass 26.8% of legal tasks, 54.6% of prescription-safety cases, and 4.9% of optimization problems.

Benchmarks

3

Categories

3

Tasks & cases

1,583

Systems evaluated

9

Browse benchmarks by category

LegalIndustry benchmark
Updated · 7/22/2026

Harvey Legal Agent Benchmark

41 tasks · 6 systems

Tests an agent's ability to complete legal work using documents, spreadsheets, presentations, and file-system tools.

Top systems

  1. Pramaana100.0%
  2. Muse Spark 1.1 (xhigh)26.8%
  3. Claude Fable 5 (max)22.0%
View benchmark
HealthcareAcademic benchmark
Updated · 7/22/2026

RxSafeBench

1,380 cases · 5 systems

Tests whether a medical agent can prescribe safely when drugs interact - picking the medication that treats the indication without harming the patient's existing regimen.

Top systems

  1. Pramaana100.0%
  2. Claude Fable 5 (max)54.6%
  3. GPT-5.6 Sol (xhigh)52.3%
View benchmark
IndustrialIndustry benchmark
Updated · 7/22/2026

Industrial-162

162 problems · 4 systems

Tests whether an agent can formulate and solve industrial-scale combinatorial optimization problems - returning verified optimal solutions, or proofs of infeasibility, with solver tools.

Top systems

  1. Pramaana100.0%
  2. Claude Fable 5 (max)4.9%
  3. GPT-5.6 Sol (xhigh)2.5%
View benchmark

Platform

Benchmarks

Company

NewsAbout us

Legal

Privacy

Research notes, essays, and benchmarks on AI verification, formal methods, and the domains being rebuilt around machine-checkable truth.

General

hello@pramaanalabs.ai

Press

press@pramaanalabs.ai

© 2026 Pramaana Labs. All rights reserved.