Overview
Harvey's Legal Agent Benchmark (LAB) measures the performance of AI agents on tasks that lawyers perform: 1,250+ tasks across 24 practice areas, graded against 75,000+ expert-written rubric criteria. Grading is all-pass - a task counts as resolved only when every one of its criteria passes.
A third of the benchmark involves taking contract playbooks and applying them consistently across multiple turns of contract negotiation. Pramaana formalized four of those playbooks roughly 10% of all contract tasks and achieved a perfect score on every one of them.
Using the formalization as a 100% baseline, we measure frontier model performance on the same tasks with the same all-pass grading.
Key takeaways
- Pramaana's formalized playbook system scores 100% - 4 task groups and all 2,564 rubric criteria pass, across all four task groups.
- Frontier performance is jagged. Meta's Muse Spark 1.1 performs best at 26.8% of tasks passed - about 7 points above its mean score reported on vals.ai - while Grok 4.5 lands below its average at 9.8%. Performance is highly variable across task subsets.
- Long-running tasks take a toll. Frontier models routinely pass rubric criteria at around 90% or higher, yet all-pass task completion stays below 30% - task length and criteria volume cause at least one failure on most tasks.
Benchmark scope
- Follow partner-style instructions across multiple turns of contract negotiation
- Navigate a closed-universe client matter with Read, Edit, Write, Glob, Bash, and Grep tools
- Produce reviewable legal work product using docx, pptx, and xlsx skills
- Satisfy every expert-written rubric criterion - all-pass grading with no partial credit