Evaluation
Making AI-Generated Legal Reasoning Testable
Today, we are introducing a new approach to testing the reliability of AI-generated legal reasoning: measuring generated work against its underlying sources across four distinct dimensions of failure.
Intelligend8 min read
Large language models have become remarkably good at producing legal analysis.
They can read lengthy judgments, identify relevant legal principles, synthesise facts and produce answers that are fluent, structured and persuasive.
But increasingly capable models create a new problem:
How do we know whether the analysis they generate actually represents its sources correctly?
Hallucination detection addresses only part of this problem.
An AI system does not need to invent a fact or citation to produce unreliable legal work. It can omit information that changes the meaning of a conclusion. It can present disputed evidence as established fact. It can preserve individual facts while losing their legal significance. And it can reach a plausible conclusion through reasoning that is incomplete or unsupported.
The resulting analysis can look entirely correct. But the underlying reasoning is flawed.
Why accuracy is not enough for legal reasoning
Most Legal AI benchmarks are built around a relatively simple idea: there is an expected answer, and we measure whether the model produces it.
This works well for tasks where a sufficiently determinate answer exists.
Professional reasoning is different.
Two competent lawyers can analyse the same materials differently. They can emphasise different facts, construct different arguments and sometimes reach different interpretations without either analysis necessarily being erroneous.
This makes a single notion of “accuracy” insufficient for evaluating open-ended legal reasoning. The problem becomes particularly important as AI moves towards generating substantive professional work.
The central question is: given the information available to the AI, does the work it produces represent that information correctly — or does it fail in ways that make its reasoning unreliable?
That question became the starting point for several years of our research.
Four ways AI-generated legal reasoning can fail
We manually analysed hundreds of AI-generated explanations of court decisions, focusing particularly on outputs that appeared plausible but failed to faithfully represent the underlying legal reasoning. Four recurring types of failure emerged.
- 01
Misleading
The information presented may be individually accurate, but the resulting account creates a misleading representation of the underlying case.
- 02
Incomplete
Legally significant information contained in the source is missing from the generated analysis.
- 03
False
The output contains factual or legal representations inconsistent with the underlying source.
- 04
Deficient
There are failures in the reasoning connecting facts, legal principles and conclusions.
Making these failures measurable
Identifying these patterns manually was the first step.
The more difficult question was whether they could be tested systematically and automatically. Over the last several years, we have been developing an evaluation approach designed to do exactly that.
Rather than reducing an AI-generated legal analysis to a single accuracy score, we test its reliability across the four dimensions separately.
The objective is not to determine whether an AI has reproduced one preferred interpretation of a legal problem. It is to identify measurable failures in the work it has generated relative to the sources on which that work is based.
We have now applied the latest version of this test to the current generation of frontier AI models.
Testing today’s frontier models
We evaluated 10 frontier models across approximately 135 complex legal cases per model. The results show significant progress in frontier AI but also reveal substantial differences in how individual models fail.