
A legal-only benchmark designed and graded by qualified attorneys across various legal task types
We designed and administered a structured benchmark to measure how leading large language models perform on the work of lawyers and legal professionals. The study was loosely based on the GDPVAL evaluation framework and adapted for the specific requirements of legal practice, where reasoning, citation accuracy, and faithful adherence to controlling authority matter.
The benchmark exercise included legal task types across various practice areas: litigation, contract redlining, employment, and insurance coverage. Each task asks the model to produce an actual deliverable a lawyer would deliver to a client or senior partner: a memo, a coverage opinion, a redlined contract, or a coded document review set with a privilege log.
Outputs from frontier models across Anthropic, OpenAI, Google, xAI, Moonshot AI, and DeepSeek were collected and blindly graded by qualified legal professionals using detailed, pre-determined rubrics. The results show meaningful and consistent differentiation between models, with reasoning-focused and extended-thinking variants outperforming standard configurations on the most demanding analytical work.
This article describes the study design, methodology, scoring framework, models evaluated, and the results.
Why a Legal AI Benchmark?
Many public benchmarks for artificial intelligence models may not optimally grade how AI performs on legal work. This is so because the LLM benchmarking exercises are often combined with other disciplines, are self graded or rely on bar-exam style questions.
While these benchmarking exercises are thoughtful and valuable, the nuance of true legal work is hard to master and measure. A practicing lawyer is not asked to choose between four multiple choice options. Lawyers are asked to read a complaint, an insurance policy or employment file, identify the issues that matter and apply the controlling authority to produce a deliverable that a client or senior partner may rely on. Whether a model can do this, and how reliably, is the question Percipient set out to answer.
The benchmark design was based on four pillars:
- Real-world legal issues. Every task is grounded in a realistic fact pattern, with reference materials a practicing lawyer would actually use: insurance policies, contract playbooks, employment files, and binding case law. Material for each task also includes deliberate noise: facts or documents that look topically relevant but are not germane to the legal task. This tests whether a model is applying genuine legal judgment or pattern-matching on surface features.
- Multimodal deliverables. For the benchmarking exercise. the models are not asked to just summarize legal documents or provide high level legal analysis. The benchmark prompts require an output a lawyer would actually deliver–a coverage opinion, an employment law memo, a coded document review with privilege log, or a redlined contract with comments. Accurate format is part of the score, because in practice a messy deliverable has diminished value.
- Blinded, rubric-based grading. Rubrics were created by legal professionals and finalized before output review. Graders scored each deliverable without knowing which model produced it. Where two graders scored the same set, their passes are kept independent and reconciled only after both were complete.
- Qualified legal professionals. All model output and deliverables were graded by attorneys with an average of over 25 years of legal experience. Domain experts were used for each corresponding area of law.
The Legal Tasks
Each task includes a structured prompt, supporting reference materials, and a defined deliverable format. The legal tasks are summarized below.
.percipient-tasks-table { width: 100%; border-collapse: collapse; font-size: 14px; }
.percipient-tasks-table thead tr { background-color: #3f6600; }
.percipient-tasks-table thead th { color: #ffffff; padding: 10px 12px; text-align: left; border: 1px solid #2d4d00; }
.percipient-tasks-table tbody tr:nth-child(odd) { background-color: #ffffff; }
.percipient-tasks-table tbody tr:nth-child(even) { background-color: #f2f2f2; }
.percipient-tasks-table tbody td { padding: 10px 12px; vertical-align: top; border: 1px solid #cccccc; }
.percipient-tasks-wrapper { overflow-x: auto; -webkit-overflow-scrolling: touch; }
@media (max-width: 768px) {
.percipient-tasks-table { font-size: 12px; min-width: 600px; }
.percipient-tasks-table thead th,
.percipient-tasks-table tbody td { padding: 8px; }
}
| Task | Description | Primary Grading Dimensions |
|---|---|---|
| Insurance Coverage Memo | Analyze a duty-to-defend scenario under Illinois law involving a construction defect / water-infiltration suit. Produce a coverage memorandum citing the eight-corners rule, the Illinois Supreme Court’s M/I Homes occurrence analysis, the subcontractor exception, and reservation-of-rights doctrine. | Eight-corners application; occurrence analysis; subcontractor cases; business-risk exclusions; reservation of rights; legal accuracy and absence of hallucinations; format and writing |
| Employment Law Memo | Analyze a former employee’s potential discrimination claims under the ADA, Title VII, and the ADEA, including a circuit split on reassignment as accommodation. Recommend a motion-stage posture and a settlement strategy responsive to the client’s framing. | Procedural and jurisdictional issues; ADA accommodation analysis; termination claims; motion-stage and trial likelihood; settlement strategy; quality control and noise resistance; format and writing |
| Litigation Document Review | Code 78 simulated litigation documents for responsiveness, privilege, issue tags, and Hot status against a pre-coded gold standard. Produce a privilege log for privileged items. | Coding accuracy across responsiveness, privilege, issue tags, and Hot calls; reasoning quality; privilege log format and sufficiency |
| Contract Review and Redline | Review a vendor agreement against a defined Vendor Contract Playbook. Identify each clause that deviates from playbook position, generate redline language, and provide reasoning for each change. Deliver a redlined .docx with comments. | Edit completeness; edit precision; reasoning aligned with playbook; absence of hallucinations; formatting of redlined deliverable and comments |
Models Evaluated
The evaluation measured models from Frontier AI companies. As each task was administered, the then-current production version of each model was used. Where extended-thinking or reasoning variants were available, they were included alongside their standard counterparts.
The employment and contract review benchmarks, which were administered after the coverage and document review work, used updated model versions. The master dataset preserves these distinctions in a private key so that grading remains blind by memo number and is not biased by knowledge of which specific model produced a given output.
.percipient-providers-table { width: 100%; border-collapse: collapse; font-size: 14px; }
.percipient-providers-table thead tr { background-color: #3f6600; }
.percipient-providers-table thead th { color: #ffffff; padding: 10px 12px; text-align: left; border: 1px solid #2d4d00; }
.percipient-providers-table tbody tr:nth-child(odd) { background-color: #ffffff; }
.percipient-providers-table tbody tr:nth-child(even) { background-color: #f2f2f2; }
.percipient-providers-table tbody td { padding: 10px 12px; vertical-align: top; border: 1px solid #cccccc; }
.percipient-providers-wrapper { overflow-x: auto; -webkit-overflow-scrolling: touch; }
@media (max-width: 768px) {
.percipient-providers-table { font-size: 12px; min-width: 480px; }
.percipient-providers-table thead th,
.percipient-providers-table tbody td { padding: 8px; }
}
| Provider | Model Family | Variants Evaluated |
|---|---|---|
| Anthropic | Claude | Opus 4.6, Opus 4.6 (extended), Sonnet 4.6, Opus 4.7 (Adaptive) |
| OpenAI | GPT / ChatGPT | 5.4 Pro (standard), 5.4 Pro (extended), 5.4 Thinking (extended), GPT 5.5 Thinking |
| Gemini | 3.0 Pro, 3.1 Pro | |
| xAI | Grok | 4.20, 4.3 (beta) |
| Moonshot AI | Kimi | K2.5, K2.5 Thinking, K2.6 Thinking |
| DeepSeek | DeepSeek | Standard |
All models received identical prompts with no model-specific tuning. No fine-tuned or custom-deployed versions were used.
Methodology
Evaluation methodology followed a structured workflow designed to minimize bias at every stage.
.percipient-phases-table { width: 100%; border-collapse: collapse; font-size: 14px; }
.percipient-phases-table thead tr { background-color: #3f6600; }
.percipient-phases-table thead th { color: #ffffff; padding: 10px 12px; text-align: left; border: 1px solid #2d4d00; }
.percipient-phases-table tbody tr:nth-child(odd) { background-color: #ffffff; }
.percipient-phases-table tbody tr:nth-child(even) { background-color: #f2f2f2; }
.percipient-phases-table tbody td { padding: 10px 12px; vertical-align: top; border: 1px solid #cccccc; }
.percipient-phases-wrapper { overflow-x: auto; -webkit-overflow-scrolling: touch; }
@media (max-width: 768px) {
.percipient-phases-table { font-size: 12px; min-width: 420px; }
.percipient-phases-table thead th,
.percipient-phases-table tbody td { padding: 8px; }
}
| Phase | Description |
|---|---|
| Prompt drafting | A two-person team takes ownership of each task. Person 1 drafts the prompt, supporting reference materials, and the deliverable specification. |
| Peer review | Person 2 reviews the prompt for clarity, completeness, and any ambiguity that might skew model outputs. Comments are returned to Person 1. |
| Prompt finalization | Person 1 incorporates feedback and finalizes the prompt that will be used uniformly across all model evaluations. |
| Rubric development | Person 2 independently develops the scoring rubric. The rubric is finalized before any model output — or any human deliverable — is reviewed. |
| Blind model evaluation | Percipient runs the prompt through each model in the roster and compiles all outputs in anonymized form. |
| Independent scoring | Graders score all deliverables using the rubric, blind to authorship. Where two graders score the same set, passes are kept independent. Final scores reflect consensus or majority position with dissent noted. |
Grading Framework
Each rubric was task-specific and used partial-credit scoring at the sub-task level. Sub-task grading roll up into section subtotals, and section subtotals roll up into a 100-point overall total. Rubrics were anchored to specific authority and specific facts in each prompt so that grading reflects whether the model did the legal work the task required, not whether the writing was generally sound.
Coverage Rubric
.percipient-rubric-ins-table { width: 100%; border-collapse: collapse; font-size: 14px; }
.percipient-rubric-ins-table thead tr { background-color: #3f6600; }
.percipient-rubric-ins-table thead th { color: #ffffff; padding: 10px 12px; text-align: left; border: 1px solid #2d4d00; }
.percipient-rubric-ins-table tbody tr:nth-child(odd) { background-color: #ffffff; }
.percipient-rubric-ins-table tbody tr:nth-child(even) { background-color: #f2f2f2; }
.percipient-rubric-ins-table tbody td { padding: 10px 12px; vertical-align: top; border: 1px solid #cccccc; }
.percipient-rubric-ins-wrapper { overflow-x: auto; -webkit-overflow-scrolling: touch; }
@media (max-width: 768px) {
.percipient-rubric-ins-table { font-size: 12px; min-width: 480px; }
.percipient-rubric-ins-table thead th,
.percipient-rubric-ins-table tbody td { padding: 8px; }
}
| Section | Max | Representative Criteria |
|---|---|---|
| Duty to Defend / Eight Corners | 15 | Eight-corners rule; scope of duty; potential coverage standard |
| Occurrence / M/I Homes | 10 | Acuity v. M/I Homes; own work vs. other property; complaint scope |
| Property Damage / Subcontractor Cases | 15 | JP Larsen; West Van Buren / Metropolitan Builders; 950 W. Huron |
| Business Risk Exclusions | 15 | Coverage A exclusions (j)/(k)/(l)/(m); subcontractor exception; PCOH |
| Reservation of Rights | 15 | Duty-to-defend conclusion; RoR analysis; Ehlco estoppel and DJ option |
| Legal Accuracy | 15 | Hallucinations; irrelevant or incorrect arguments |
| Format, Professionalism, Writing | 15 | Memo format; conservative tone; balanced authority; citations |
Employment Rubric
.percipient-rubric-emp-table { width: 100%; border-collapse: collapse; font-size: 14px; }
.percipient-rubric-emp-table thead tr { background-color: #3f6600; }
.percipient-rubric-emp-table thead th { color: #ffffff; padding: 10px 12px; text-align: left; border: 1px solid #2d4d00; }
.percipient-rubric-emp-table tbody tr:nth-child(odd) { background-color: #ffffff; }
.percipient-rubric-emp-table tbody tr:nth-child(even) { background-color: #f2f2f2; }
.percipient-rubric-emp-table tbody td { padding: 10px 12px; vertical-align: top; border: 1px solid #cccccc; }
.percipient-rubric-emp-wrapper { overflow-x: auto; -webkit-overflow-scrolling: touch; }
@media (max-width: 768px) {
.percipient-rubric-emp-table { font-size: 12px; min-width: 480px; }
.percipient-rubric-emp-table thead th,
.percipient-rubric-emp-table tbody td { padding: 8px; }
}
| Section | Max | Representative Criteria |
|---|---|---|
| Procedural and Jurisdictional | 10 | Administrative exhaustion; state law claims; circuit analysis |
| ADA Failure to Accommodate | 20 | Prima facie framework; qualified individual; reassignment / EEOC v. UAL; interactive process; Dark v. Curry County noise test |
| Termination Claims | 15 | ADA termination; ADA retaliation; Title VII race and gender; ADEA |
| Motion Stage and Trial Likelihood | 10 | Motion to dismiss; summary judgment; trial outcome assessment |
| Settlement Strategy and Risk | 15 | Direct recommendation; damages exposure; cost of litigation; intangible costs; negotiating posture |
| Quality Control | 15 | Hallucinations; irrelevant arguments and noise resistance; assumptions and missing information |
| Format, Professionalism, Writing | 15 | Memo format; outline / bullet structure; Bluebook citations; tone; organization |
Document Review Rubric
.percipient-rubric-lit-table { width: 100%; border-collapse: collapse; font-size: 14px; }
.percipient-rubric-lit-table thead tr { background-color: #3f6600; }
.percipient-rubric-lit-table thead th { color: #ffffff; padding: 10px 12px; text-align: left; border: 1px solid #2d4d00; }
.percipient-rubric-lit-table tbody tr:nth-child(odd) { background-color: #ffffff; }
.percipient-rubric-lit-table tbody tr:nth-child(even) { background-color: #f2f2f2; }
.percipient-rubric-lit-table tbody td { padding: 10px 12px; vertical-align: top; border: 1px solid #cccccc; }
.percipient-rubric-lit-wrapper { overflow-x: auto; -webkit-overflow-scrolling: touch; }
@media (max-width: 768px) {
.percipient-rubric-lit-table { font-size: 12px; min-width: 380px; }
.percipient-rubric-lit-table thead th,
.percipient-rubric-lit-table tbody td { padding: 8px; }
}
| Section | Max | Representative Criteria |
|---|---|---|
| Coding | 80 | Thoroughness — issues identified; completeness — accuracy; reasoning; redlined .docx; comments |
| Privilege Log | 20 | Format; reasoning; sufficiency |
Contract Review Rubric
.percipient-rubric-contract-table { width: 100%; border-collapse: collapse; font-size: 14px; }
.percipient-rubric-contract-table thead tr { background-color: #3f6600; }
.percipient-rubric-contract-table thead th { color: #ffffff; padding: 10px 12px; text-align: left; border: 1px solid #2d4d00; }
.percipient-rubric-contract-table tbody tr:nth-child(odd) { background-color: #ffffff; }
.percipient-rubric-contract-table tbody tr:nth-child(even) { background-color: #f2f2f2; }
.percipient-rubric-contract-table tbody td { padding: 10px 12px; vertical-align: top; border: 1px solid #cccccc; }
.percipient-rubric-contract-wrapper { overflow-x: auto; -webkit-overflow-scrolling: touch; }
@media (max-width: 768px) {
.percipient-rubric-contract-table { font-size: 12px; min-width: 480px; }
.percipient-rubric-contract-table thead th,
.percipient-rubric-contract-table tbody td { padding: 8px; }
}
| Criterion | Max | What is Evaluated |
|---|---|---|
| Edit Complete | 20 | Whether the redline captures every change the playbook requires for each clause |
| Edit Precise | 20 | Whether the redline language is precisely worded and tracks the playbook position |
| Reasoning Aligned with Playbook | 20 | Whether the stated reasoning for each redline matches the playbook rationale |
| No Hallucination | 20 | Whether the output is free of invented contractual terms, authorities, or facts |
| Format — Redlined .docx | 10 | Whether a usable redlined .docx was generated |
| Format — Inline Comments | 10 | Whether inline comments were generated and free of hallucinations |
Model Results
Coverage
.percipient-results-table { width: 100%; border-collapse: collapse; font-size: 14px; }
.percipient-results-table thead tr { background-color: #3f6600; }
.percipient-results-table thead th { color: #ffffff; padding: 10px 12px; text-align: left; border: 1px solid #2d4d00; }
.percipient-results-table tbody tr:nth-child(odd) { background-color: #ffffff; }
.percipient-results-table tbody tr:nth-child(even) { background-color: #f2f2f2; }
.percipient-results-table tbody td { padding: 10px 12px; vertical-align: top; border: 1px solid #cccccc; }
.percipient-results-table tbody td:nth-child(2),
.percipient-results-table tbody td:nth-child(3),
.percipient-results-table tbody td:nth-child(4),
.percipient-results-table tbody td:nth-child(5) { text-align: center; }
.percipient-results-wrapper { overflow-x: auto; -webkit-overflow-scrolling: touch; }
@media (max-width: 768px) {
.percipient-results-table { font-size: 12px; min-width: 520px; }
.percipient-results-table thead th,
.percipient-results-table tbody td { padding: 8px; }
}
| Model | Expert 1 / 100 | Expert 2 / 100 | Average / 100 | Rank |
|---|---|---|---|---|
| Claude Opus 4.7 (Adaptive) | 79.0 | 87.5 | 83.25 | 1 |
| ChatGPT 5.4 Pro | 76.0 | 84.5 | 80.25 | 2 |
| Grok 4.20 | 70.5 | 76.5 | 73.50 | 3 |
| Kimi K2.5 Thinking | 47.5 | 71.0 | 59.25 | 4 |
| DeepSeek | 58.5 | 53.5 | 56.00 | 5 |
| Gemini 3.0 Pro | 34.5 | 57.5 | 46.00 | 6 |
Document Review
.percipient-docreview-table { width: 100%; border-collapse: collapse; font-size: 14px; }
.percipient-docreview-table thead tr { background-color: #3f6600; }
.percipient-docreview-table thead th { color: #ffffff; padding: 10px 12px; text-align: left; border: 1px solid #2d4d00; }
.percipient-docreview-table tbody tr:nth-child(odd) { background-color: #ffffff; }
.percipient-docreview-table tbody tr:nth-child(even) { background-color: #f2f2f2; }
.percipient-docreview-table tbody td { padding: 10px 12px; vertical-align: top; border: 1px solid #cccccc; }
.percipient-docreview-table tbody td:nth-child(2),
.percipient-docreview-table tbody td:nth-child(3),
.percipient-docreview-table tbody td:nth-child(4) { text-align: center; }
.percipient-docreview-wrapper { overflow-x: auto; -webkit-overflow-scrolling: touch; }
@media (max-width: 768px) {
.percipient-docreview-table { font-size: 12px; min-width: 520px; }
.percipient-docreview-table thead th,
.percipient-docreview-table tbody td { padding: 8px; }
}
| Model | Coding / 80 | Priv. Log / 20 | Total / 100 |
|---|---|---|---|
| Claude Opus 4.6 (extended) | 77.8 | 19.0 | 96.8 |
| Claude Sonnet 4.6 | 75.5 | 18.5 | 94.0 |
| Claude Opus 4.6 | 75.4 | 18.5 | 93.9 |
| ChatGPT 5.4 Pro (extended) | 76.0 | 17.5 | 93.5 |
| ChatGPT 5.4 Pro (standard) | 74.7 | 18.5 | 93.2 |
| DeepSeek | 75.9 | 16.5 | 92.4 |
| ChatGPT 5.4 Thinking (extended) | 75.3 | 16.5 | 91.8 |
| Gemini 3.0 Pro | 73.7 | 17.0 | 90.7 |
| Kimi K2.5 | 74.4 | 15.0 | 89.4 |
| Grok 4.20 | 52.6 | 11.0 | 63.6 |
Employment
.percipient-emp-scores-table { width: 100%; border-collapse: collapse; font-size: 14px; }
.percipient-emp-scores-table thead tr { background-color: #3f6600; }
.percipient-emp-scores-table thead th { color: #ffffff; padding: 10px 12px; text-align: center; border: 1px solid #2d4d00; }
.percipient-emp-scores-table thead th:first-child { text-align: left; }
.percipient-emp-scores-table tbody tr:nth-child(odd) { background-color: #ffffff; }
.percipient-emp-scores-table tbody tr:nth-child(even) { background-color: #f2f2f2; }
.percipient-emp-scores-table tbody td { padding: 10px 12px; vertical-align: top; border: 1px solid #cccccc; text-align: center; }
.percipient-emp-scores-table tbody td:first-child { text-align: left; }
.percipient-emp-scores-wrapper { overflow-x: auto; -webkit-overflow-scrolling: touch; }
@media (max-width: 768px) {
.percipient-emp-scores-table { font-size: 12px; min-width: 720px; }
.percipient-emp-scores-table thead th,
.percipient-emp-scores-table tbody td { padding: 8px; }
}
| Model | Procedural / 10 | ADA Accomm. / 20 | Termination / 15 | Motion & Trial / 10 | Settlement / 15 | Quality Control / 15 | Format / 15 | Total / 100 |
|---|---|---|---|---|---|---|---|---|
| Kimi K2.6 Thinking | 5.0 | 15.5 | 5.5 | 5.0 | 8.0 | 9.0 | 8.5 | 56.5 |
| Claude Opus 4.7 (Adaptive) | 8.0 | 7.5 | 6.0 | 8.0 | 13.0 | 9.0 | 10.0 | 61.5 |
| DeepSeek | 7.0 | 9.0 | 9.5 | 8.0 | 11.5 | 8.0 | 5.0 | 58.0 |
| GPT 5.5 Thinking | 7.0 | 11.0 | 13.5 | 5.0 | 7.5 | 9.0 | 8.5 | 61.5 |
| Gemini 3.1 Pro | 6.0 | 1.5 | 3.0 | 2.5 | 10.0 | 9.0 | 9.0 | 41.0 |
| Grok 4.3 (beta) | 5.0 | 7.5 | 8.0 | 2.5 | 10.5 | 8.0 | 10.0 | 51.5 |
Contract Review
.percipient-emp-results-table { width: 100%; border-collapse: collapse; font-size: 14px; }
.percipient-emp-results-table thead tr { background-color: #3f6600; }
.percipient-emp-results-table thead th { color: #ffffff; padding: 10px 12px; text-align: left; border: 1px solid #2d4d00; }
.percipient-emp-results-table tbody tr:nth-child(odd) { background-color: #ffffff; }
.percipient-emp-results-table tbody tr:nth-child(even) { background-color: #f2f2f2; }
.percipient-emp-results-table tbody td { padding: 10px 12px; vertical-align: top; border: 1px solid #cccccc; }
.percipient-emp-results-table tbody td:nth-child(2),
.percipient-emp-results-table tbody td:nth-child(3),
.percipient-emp-results-table tbody td:nth-child(4),
.percipient-emp-results-table tbody td:nth-child(5) { text-align: center; }
.percipient-emp-results-wrapper { overflow-x: auto; -webkit-overflow-scrolling: touch; }
@media (max-width: 768px) {
.percipient-emp-results-table { font-size: 12px; min-width: 520px; }
.percipient-emp-results-table thead th,
.percipient-emp-results-table tbody td { padding: 8px; }
}
| Model | Expert 1 / 100 | Expert 2 / 100 | Average / 100 | Rank |
|---|---|---|---|---|
| Claude Opus 4.7 (Adaptive) | 88.00 | 90.25 | 89.13 | 1 |
| GPT 5.5 Thinking | 88.00 | 87.75 | 87.88 | 2 |
| DeepSeek | 86.50 | 75.25 | 80.88 | 3 |
| Grok 4.3 (beta) | 77.50 | 81.00 | 79.25 | 4 |
| Kimi K2.6 Thinking | 86.00 | 69.00 | 77.50 | 5 |
| Gemini 3.1 Pro | 80.50 | 68.00 | 74.25 | 6 |
The completed tasks support several preliminary observations.
- Models cluster tightly on routine work and deviate on analytical work. On document review, nine of ten model variants scored between 89 and 97 out of 100. On the coverage memo, the same model families spread across a 37-point range. The harder the task, the more the rubric distinguishes between systems that genuinely apply controlling authority and systems that merely sound legally fluent.
- Reasoning-mode models lead the analytical tasks. Across the coverage, employment, and contract benchmarks, tasks that require sustained multi-step legal reasoning, the top scorers were consistently extended-thinking or reasoning-mode variants. Claude Opus 4.7 (Adaptive) and GPT 5.5 Thinking were the top two on contract review and tied for the top score on employment, and Claude Opus 4.7 (Adaptive) also led the coverage benchmark.
- Extended thinking helps where the analysis is hard, not where the work is large. The extended-thinking variant of Claude Opus 4.6 was the top scorer on document review, but standard variants of the same family were within roughly three points. On the harder analytical tasks, the gap opens up and reasoning-mode variants pulled away from their standard counterparts.
- Format and writing are not free points. Across both expert passes on the coverage benchmark, the lowest-scoring model lost roughly half of its points in the legal-accuracy and writing sections combined — driven by hallucinated authority, irrelevant arguments, and a deliverable that did not satisfy the memo specification. On the contract task, format alone is worth twenty points, and the lowest format scores came from models that produced substantive analysis but failed to generate a usable redlined .docx or properly attached inline comments.
- Two-grader passes are worth the cost. On the coverage benchmark, the average difference between Expert 1 and Expert 2 scores was about 11 points; on contract review, the average gap was about 6 points but several individual outputs differed by more than 15 points. Both pairs of passes nonetheless agreed on the top of the field, suggesting the rubrics are doing real work — but it also suggests that a single-grader pass would have produced a less reliable ordering in the middle of the field.
- Noise resistance is a real differentiator. Each task includes deliberate noise, facts and documents that look topically relevant but are not legally responsive. On the employment memo, where the fact pattern includes both hallucination traps and a Dark v. Curry County noise test, the spread between the highest and lowest scorers on the Quality Control section was meaningful, even though the rubric weighting on that section is only fifteen points.
How These Results Should Be Used
This benchmark is not a vendor scorecard. It does not declare a winner. In the real world, models change, prompts vary, and the same model can perform differently on fact patterns. The point of the work is methodological: to show what a defensible legal AI evaluation looks like, to provide a repeatable framework for measuring real legal performance, and to give legal teams a basis for asking sharper questions when they evaluate a tool.
Ultimately, the most critical takeaway from this benchmark isn’t the final score, but the fact that veteran attorneys with decades of domain expertise still varied when grading complex legal problems. This proves that true legal work is deeply nuanced, subjective, and there is still much room for improvement with model outputs.
The full data set may be found on Hugging Face.
The post How Frontier AI Models Perform on Real Legal Work appeared first on Percipient – Legal Services Powered by Technology.