Real-world benchmarks for long-horizon agent work.
Three benchmarks across legal, healthcare, and industrial optimization, each measured against Pramaana's formalized baseline. Pramaana scores 100% on every benchmark - the best frontier models pass 26.8% of legal tasks, 54.6% of prescription-safety cases, and 4.9% of optimization problems.
Benchmarks
3
Categories
3
Tasks & cases
1,583
Systems evaluated
9