Legal Benchmark
How the HAQQ Legal Benchmark is built, what it measures, what it deliberately does not, and where to find the numbers it produces.
What it is
Two evaluations under one name, run on different scales because they answer different questions.
| Evaluation | Scale | Shape | The question |
|---|---|---|---|
| The task benchmark | Out of 50 | 11 legal task categories across 19 systems | On the work a lawyer actually hands over, who produces the better output? |
| The cross-domain study | Out of 100 | 25 benchmarks in 5 domains | Against the frontier models on their own ground, where does a legal-specialised engine stand? |
The two scales are not interchangeable. A score out of 50 on the task benchmark and a score out of 100 on the cross-domain study measure different work on different rubrics. Doubling one does not produce the other, and the two must never appear in the same table or the same sentence as though they were comparable.
Where the numbers live
- The practice-area leaderboard on haqq.ai: the task benchmark, grouped into ten practice areas.
- The task-level table on /justinian: the cross-domain study, every benchmark and every model, sortable and filterable.
- HAQQ Labs: both, side by side, with the research context around them.
How a score is produced
The task benchmark
Each system is given the same prompt for a task a lawyer would recognise: draft this NDA, research this question, explain this provision to a client. The output is scored against a fifty-point rubric covering Sharia, statute, forum, clause construction, risk, hallucination, formatting, brevity, partner-readiness and source linking.
Partner-readiness and source linking are the two that separate a legal evaluation from a general one. An answer that is correct but needs an hour of restructuring before a partner would sign it has not done the job, and a citation that does not resolve is worse than no citation, because it transfers the checking cost to the reader without telling them.
The cross-domain study
Twenty-five benchmarks across five domains: legal, tax, journalism, general capability, and safety and values. The legal domain draws on public frameworks rather than instruments we wrote ourselves, which is what makes the legal columns checkable by someone outside HAQQ.
| Framework | What it contributes |
|---|---|
| Stanford LegalBench | Legal reasoning across a broad task taxonomy |
| Harvey LAB | All-pass rubric grading of long-horizon legal work |
| VLAIR (Vals AI) | Independent industry evaluation of legal AI systems |
| ALARB | Arabic legal reasoning, which most legal benchmarks do not cover |
| BigLaw Bench | Tasks drawn from large-firm practice |
| CUAD | Contract clause extraction and understanding |
| LegalCiteBench | Citation accuracy and resolvability |
| LawBench / LexEval | Multilingual legal capability |
| Legal Benchmarks | Cross-vendor methodology |
We replicated the all-pass methodology on civil-law and MENA data rather than inventing a parallel instrument. All-pass means a task counts only when every criterion in its rubric is met: a near-miss scores zero, not partial credit. It is a harsh grading scheme and it is the right one for legal work, where an output that is ninety per cent correct still has to be checked line by line.
How the averages work
A domain average is the mean of that domain's benchmarks. The overall average is the mean of the twenty-four benchmarks excluding adversarial testing, because several systems were never run on that benchmark, and averaging a column that exists for some entrants and not others produces a number that is not a comparison.
An absent cell renders as blank and is excluded from every average that would otherwise contain it. It is never treated as a zero. Reading a blank as zero would move a safety average by more than twenty points, which would look like a finding and would be an artefact.
What this benchmark does not do
The limits are the part worth reading closely, because a benchmark that only advertises its coverage is advertising.
- It is not independent. HAQQ builds it, runs it, and competes in it. That is why the methodology is published and the per-task results are visible rather than summarised: the check available to you is inspection, not trust.
- It does not cover every practice area. The task benchmark is weighted toward corporate and commercial drafting, research and client-facing explanation. Areas outside that are covered thinly or not at all.
- It does not cover every system. Several legal AI vendors publish no evaluations and offer no self-serve access, so scoring them needs their cooperation. A system absent from a table was not beaten by us; it was not run.
- It is a point in time. Every entrant ships new versions. A score is the version named at the run date and nothing more.
- It measures output, not deployment. Security posture, data residency, support and integration depth decide whether a system can be used at all, and none of them appear here. Those are on Trust.
Reading a result honestly
Rows where HAQQ is not first are rendered exactly like rows where it is: same weight, same position, no reordering and no omission. If a specialist beats us on a task, the table says so, because a leaderboard that never shows its author losing is not reporting a measurement.
Where a figure is derived rather than measured directly, the derivation is stated. On the practice-area leaderboard, a practice area's score is the mean of the measured task categories mapped to it, and the mapping is published in the dataset so any cell can be checked by hand.
Related
Trust covers security posture and compliance scope. Product updates records what shipped in each release. The engine the benchmark measures is described on Justinian.