BenchmarksLegal · Harvey LAB
LegalIndustry benchmark

Harvey Legal Agent Benchmark

Updated: 7/22/2026

Tests an agent's ability to complete legal work using documents, spreadsheets, presentations, and file-system tools.

41 tasks·4 task groups·2,564 criteria·6 systems

Pramaana scores 100% across everything

Every task, every criterion, every task group - a perfect score from Pramaana's formalized playbook system.

Criteria
2,564
Groups
4
LEGAL AGENT VERIFIED · PRAMAANA LABS100%VERIFIED

Results

#SystemTask pass rateCriteria
1
PramaanaPramaana Labs
100.0% ± 0.0100.0%
2
Muse Spark 1.1 (xhigh)MetaBest frontier model
26.8% ± 6.995.7%
3
Claude Fable 5 (max)Anthropic
22.0% ± 6.592.4%
4
Claude Opus 4.8 (max)Anthropic
14.6% ± 5.592.8%
5
GPT-5.5 (xhigh)OpenAI
12.2% ± 5.186.0%
6
Grok 4.5SpaceXAI
9.8% ± 4.688.5%

Pramaana evaluation runs on the formalized subset - 41 contract tasks, 2,564 criteria, all-pass grading per Harvey's protocol. GPT-5.5 (xhigh) returned incomplete grading on 201 criteria; ungraded criteria are counted as failures.

Criteria pass rate by task group

#SystemEnergyMediaEmployment & CompensationClinical Trial Agreements
1
PramaanaPramaana Labs
100.0%100.0%100.0%100.0%
2
Muse Spark 1.1 (xhigh)Meta
93.8%92.5%96.7%97.0%
3
Claude Fable 5 (max)Anthropic
92.6%93.5%94.5%82.2%
4
Claude Opus 4.8 (max)Anthropic
91.5%91.0%93.5%93.4%
5
GPT-5.5 (xhigh)OpenAI
92.9%94.3%81.1%90.4%
6
Grok 4.5SpaceXAI
88.5%91.5%87.0%91.3%

Energy - 339 criteria · Media - 387 criteria · Employment & Compensation - 1,472 criteria · Clinical Trial Agreements - 366 criteria

Overview

Harvey's Legal Agent Benchmark (LAB) measures the performance of AI agents on tasks that lawyers perform: 1,250+ tasks across 24 practice areas, graded against 75,000+ expert-written rubric criteria. Grading is all-pass - a task counts as resolved only when every one of its criteria passes.

A third of the benchmark involves taking contract playbooks and applying them consistently across multiple turns of contract negotiation. Pramaana formalized four of those playbooks roughly 10% of all contract tasks and achieved a perfect score on every one of them.

Using the formalization as a 100% baseline, we measure frontier model performance on the same tasks with the same all-pass grading.

Open Harvey benchmark announcement

Key takeaways

  • Pramaana's formalized playbook system scores 100% - 4 task groups and all 2,564 rubric criteria pass, across all four task groups.
  • Frontier performance is jagged. Meta's Muse Spark 1.1 performs best at 26.8% of tasks passed - about 7 points above its mean score reported on vals.ai - while Grok 4.5 lands below its average at 9.8%. Performance is highly variable across task subsets.
  • Long-running tasks take a toll. Frontier models routinely pass rubric criteria at around 90% or higher, yet all-pass task completion stays below 30% - task length and criteria volume cause at least one failure on most tasks.

Benchmark scope

  • Follow partner-style instructions across multiple turns of contract negotiation
  • Navigate a closed-universe client matter with Read, Edit, Write, Glob, Bash, and Grep tools
  • Produce reviewable legal work product using docx, pptx, and xlsx skills
  • Satisfy every expert-written rubric criterion - all-pass grading with no partial credit

Updates

4 playbooks formalized

We have formalized four contract playbooks so far - Energy, Media, Employment, and Clinical Trials - covering 4 tasks groups and 2,564 rubric criteria. Pramaana passes 100% of tasks and criteria across all four groups. More categories are coming soon.