Skip to content

Legal Benchmark

How the HAQQ Legal Benchmark is built, what it measures, what it deliberately does not, and where to find the numbers it produces.

What it is

Two evaluations under one name, run on different scales because they answer different questions.

EvaluationScaleShapeThe question
The task benchmarkOut of 5011 legal task categories across 19 systemsOn the work a lawyer actually hands over, who produces the better output?
The cross-domain studyOut of 10025 benchmarks in 5 domainsAgainst the frontier models on their own ground, where does a legal-specialised engine stand?

The two scales are not interchangeable. A score out of 50 on the task benchmark and a score out of 100 on the cross-domain study measure different work on different rubrics. Doubling one does not produce the other, and the two must never appear in the same table or the same sentence as though they were comparable.

The task benchmark

Eleven measured task categories, scored out of 50. Best in each row is bold. These are the figures the practice-area leaderboard on haqq.ai groups into practice areas.

Task benchmark, out of 50, top five per category
Category1st2nd3rd4th5th
LegalHAQQ (Justinian) 49Anthropic: Claude Fable 5 45Anthropic: Claude Opus 4.8 43Mike OS 42DeepSeek: DeepSeek V4 Pro 40
Contract DraftingHAQQ (Justinian) 47Spellbook 46Anthropic: Claude Fable 5 45Anthropic: Claude Opus 4.8 43Legora 42
Legal ResearchHAQQ (Justinian) 47LexisNexis +AI 46Perplexity Sonar 43Anthropic: Claude Fable 5 41Thomson Reuters CoCounsel 40
Law ExplanationHAQQ (Justinian) 46OpenAI: GPT-5.6 Sol Pro 45Anthropic: Claude Fable 5 43Google: Gemini 3.1 Pro 41Anthropic: Claude Opus 4.8 39
Employment AgreementHAQQ (Justinian) 48Anthropic: Claude Fable 5 43Harvey Tenet 42Mike OS 41Legora 40
Memo DraftingHAQQ (Justinian) 44Thomson Reuters CoCounsel 42Anthropic: Claude Fable 5 41Anthropic: Claude Opus 4.8 41Mike OS 39
License AgreementHAQQ (Justinian) 46Anthropic: Claude Opus 4.8 43Anthropic: Claude Fable 5 42Spellbook 41Legora 40
Shareholder AgreementHAQQ (Justinian) 47Harvey Tenet 44Anthropic: Claude Fable 5 42Mike OS 41Anthropic: Claude Opus 4.8 40
Consultancy AgreementHAQQ (Justinian) 45Anthropic: Claude Fable 5 42Legora 41Mike OS 40Anthropic: Claude Opus 4.8 39
Commercial AgreementHAQQ (Justinian) 47Harvey Tenet 43Anthropic: Claude Opus 4.8 42Anthropic: Claude Fable 5 41Mike OS 40
NDA DraftingHAQQ (Justinian) 49Anthropic: Claude Fable 5 45Spellbook 44Mike OS 42Anthropic: Claude Opus 4.8 41

The cross-domain study

Twenty-five benchmarks in five domains, scored out of 100. The same run the task-level table on /justinian sorts and filters.

Cross-domain study, out of 100
BenchmarkHAQQ (Justinian)Anthropic: Claude Opus 4.8Thomson Reuters CoCounselGoogle: Gemini 3.1 ProZ.ai: GLM 5.2OpenAI: GPT-5.6 Sol ProKimi: K3Anthropic: Claude Sonnet 5Snowdon 1.0-LargeQwen: Qwen3.7 PlusDeepSeek: DeepSeek V4 Pro
Overall average83.279.578.578.077.276.575.775.773.673.071.9
Stanford LegalBench85.781.882.384.382.982.383.381.482.878.876.8
Info. Retrieval57.251.253.653.853.653.853.047.951.249.051.5
Reasoning78.275.273.277.873.168.474.873.170.866.570.1
Classification74.770.570.974.270.270.071.670.469.868.767.8
Doc. Processing & RAG83.278.979.778.276.578.679.074.771.274.173.7
Summarisation90.088.390.089.490.086.387.784.887.982.589.1
Contract Under.77.474.471.075.973.674.977.269.671.768.468.3
Human Queries89.785.489.286.487.388.661.484.387.186.786.4
Deep Research90.990.888.980.689.085.989.886.078.487.384.0
Harvey LAB86.886.985.755.584.676.183.780.956.370.683.1
Legal average81.478.378.475.678.176.576.175.372.773.375.1
Deep Research86.385.882.476.083.580.985.182.070.478.780.7
Tax Q&A88.986.987.984.585.786.688.388.788.285.182.8
Tax average87.686.385.180.284.683.886.785.479.381.981.8
Deep Research84.682.380.983.079.066.384.574.176.277.278.6
Journalism average84.682.380.983.079.066.384.574.176.277.278.6
Factuality82.971.473.782.666.268.369.559.871.776.269.1
Long Context75.675.275.375.075.970.073.570.774.653.069.5
Multilingualism85.883.278.485.781.782.784.979.574.277.678.4
Instr. Following91.786.191.484.889.489.085.886.887.385.689.2
Writing80.779.380.378.578.179.178.278.779.977.980.7
Reasoning76.273.768.474.867.371.676.266.866.866.665.0
General Agent89.283.489.184.477.575.661.074.187.185.572.0
Coding66.857.439.950.056.050.866.857.440.943.945.6
Maths98.798.794.097.491.297.796.285.995.594.994.9
General average83.178.776.779.275.976.176.973.375.373.573.8
Political Neutrality97.982.897.393.391.083.833.582.385.351.531.3
Robustness78.578.260.365.350.468.372.777.142.266.738.1
Adversarial Testing95.993.493.687.378.895.981.1
Safety / Values average90.883.678.364.568.771.450.2

An empty cell means the model was not run on that benchmark. It is not a zero, and it is excluded from that model's averages.

Where else these numbers appear

All four surfaces read one dataset, so they cannot disagree about a run.

How a score is produced

The task benchmark

Each system is given the same prompt for a task a lawyer would recognise: draft this NDA, research this question, explain this provision to a client. The output is scored against a fifty-point rubric covering Sharia, statute, forum, clause construction, risk, hallucination, formatting, brevity, partner-readiness and source linking.

Partner-readiness and source linking are the two that separate a legal evaluation from a general one. An answer that is correct but needs an hour of restructuring before a partner would sign it has not done the job, and a citation that does not resolve is worse than no citation, because it transfers the checking cost to the reader without telling them.

The cross-domain study

Twenty-five benchmarks across five domains: legal, tax, journalism, general capability, and safety and values. The legal domain draws on public frameworks rather than instruments we wrote ourselves, which is what makes the legal columns checkable by someone outside HAQQ.

Public frameworks this study draws on, each linked to its source
FrameworkWhat it contributes
Stanford LegalBenchLegal reasoning across a broad task taxonomy
Harvey LABAll-pass rubric grading of long-horizon legal work
VLAIR (Vals AI)Independent industry evaluation of legal AI systems
ALARBArabic legal reasoning, which most legal benchmarks do not cover
BigLaw BenchTasks drawn from large-firm practice
CUADContract clause extraction and understanding
LegalCiteBenchCitation accuracy and resolvability
LawBench / LexEvalMultilingual legal capability
Legal BenchmarksCross-vendor methodology

We replicated the all-pass methodology on civil-law and MENA data rather than inventing a parallel instrument. All-pass means a task counts only when every criterion in its rubric is met: a near-miss scores zero, not partial credit. It is a harsh grading scheme and it is the right one for legal work, where an output that is ninety per cent correct still has to be checked line by line.

How the averages work

A domain average is the mean of that domain's benchmarks. The overall average is the mean of the twenty-four benchmarks excluding adversarial testing, because several systems were never run on that benchmark, and averaging a column that exists for some entrants and not others produces a number that is not a comparison.

An absent cell renders as blank and is excluded from every average that would otherwise contain it. It is never treated as a zero. Reading a blank as zero would move a safety average by more than twenty points, which would look like a finding and would be an artefact.

What this benchmark does not do

The limits are the part worth reading closely, because a benchmark that only advertises its coverage is advertising.

  • It is not independent. HAQQ builds it, runs it, and competes in it. That is why the methodology is published and the per-task results are visible rather than summarised: the check available to you is inspection, not trust.
  • It does not cover every practice area. The task benchmark is weighted toward corporate and commercial drafting, research and client-facing explanation. Areas outside that are covered thinly or not at all.
  • It does not cover every system. Several legal AI vendors publish no evaluations and offer no self-serve access, so scoring them needs their cooperation. A system absent from a table was not beaten by us; it was not run.
  • It is a point in time. Every entrant ships new versions. A score is the version named at the run date and nothing more.
  • It measures output, not deployment. Security posture, data residency, support and integration depth decide whether a system can be used at all, and none of them appear here. Those are on Trust.

Reading a result honestly

Rows where HAQQ is not first are rendered exactly like rows where it is: same weight, same position, no reordering and no omission. If a specialist beats us on a task, the table says so, because a leaderboard that never shows its author losing is not reporting a measurement.

Where a figure is derived rather than measured directly, the derivation is stated. On the practice-area leaderboard, a practice area's score is the mean of the measured task categories mapped to it, and the mapping is published in the dataset so any cell can be checked by hand.

The model router

A benchmark is only useful if it changes what you actually run. The legal AI model router is the open-source companion to this work: a vendor-neutral set of skills that recommends which model to use for a given legal task, grounded in the same published benchmarks rather than in a vendor's preference.

It is deliberately not a HAQQ recommendation engine. It is AGPL-3.0, it names other vendors' models where they win, and its scorecard is a single file anyone can read and argue with. A router that always returned our own engine would be a price list.

What the router covers
SkillThe question it answers
legal-ai-model-routerWhich of the five below applies to this task?
route-contract-draftingWhich model drafts this instrument best?
route-contract-reviewWhich model reviews and redlines it best?
route-legal-researchWhich model researches this question best?
route-info-extractionWhich model extracts structured data from these documents?
route-legal-translationWhich model translates this legal text best?

Trust covers security posture and compliance scope. Product updates records what shipped in each release. The engine the benchmark measures is described on Justinian.

Argomenti correlati

Questa pagina ti è stata utile?