From the lab

Evaluation

Making AI-Generated Legal Reasoning Testable

Today, we are introducing a new approach to testing the reliability of AI-generated legal reasoning: measuring generated work against its underlying sources across four distinct dimensions of failure.

Intelligend8 min read

Large language models have become remarkably good at producing legal analysis.

They can read lengthy judgments, identify relevant legal principles, synthesise facts and produce answers that are fluent, structured and persuasive.

But increasingly capable models create a new problem:

How do we know whether the analysis they generate actually represents its sources correctly?

Hallucination detection addresses only part of this problem.

An AI system does not need to invent a fact or citation to produce unreliable legal work. It can omit information that changes the meaning of a conclusion. It can present disputed evidence as established fact. It can preserve individual facts while losing their legal significance. And it can reach a plausible conclusion through reasoning that is incomplete or unsupported.

The resulting analysis can look entirely correct. But the underlying reasoning is flawed.

Why accuracy is not enough for legal reasoning

Most Legal AI benchmarks are built around a relatively simple idea: there is an expected answer, and we measure whether the model produces it.

This works well for tasks where a sufficiently determinate answer exists.

Professional reasoning is different.

Two competent lawyers can analyse the same materials differently. They can emphasise different facts, construct different arguments and sometimes reach different interpretations without either analysis necessarily being erroneous.

This makes a single notion of “accuracy” insufficient for evaluating open-ended legal reasoning. The problem becomes particularly important as AI moves towards generating substantive professional work.

The central question is: given the information available to the AI, does the work it produces represent that information correctly — or does it fail in ways that make its reasoning unreliable?

That question became the starting point for several years of our research.

Four ways AI-generated legal reasoning can fail

We manually analysed hundreds of AI-generated explanations of court decisions, focusing particularly on outputs that appeared plausible but failed to faithfully represent the underlying legal reasoning. Four recurring types of failure emerged.

  1. 01

    Misleading

    The information presented may be individually accurate, but the resulting account creates a misleading representation of the underlying case.

  2. 02

    Incomplete

    Legally significant information contained in the source is missing from the generated analysis.

  3. 03

    False

    The output contains factual or legal representations inconsistent with the underlying source.

  4. 04

    Deficient

    There are failures in the reasoning connecting facts, legal principles and conclusions.

Making these failures measurable

Identifying these patterns manually was the first step.

The more difficult question was whether they could be tested systematically and automatically. Over the last several years, we have been developing an evaluation approach designed to do exactly that.

Rather than reducing an AI-generated legal analysis to a single accuracy score, we test its reliability across the four dimensions separately.

The objective is not to determine whether an AI has reproduced one preferred interpretation of a legal problem. It is to identify measurable failures in the work it has generated relative to the sources on which that work is based.

We have now applied the latest version of this test to the current generation of frontier AI models.

Testing today’s frontier models

We evaluated 10 frontier models across approximately 135 complex legal cases per model. The results show significant progress in frontier AI but also reveal substantial differences in how individual models fail.

ModelMisleading divergence ↓Incomplete ↓Reasoning flaws / case ↓False statements ↓Cases with ≥1 false statement ↓
GPT-6 Sol17.0%43.5%7.340.84%12.6%
GPT-6 Astra17.8%29.4%9.080.73%21.5%
Claude Opus 5.525.9%24.3%12.931.87%67.2%
Kimi K320.7%26.8%10.612.44%62.2%
Grok 4.625.9%41.5%8.922.67%38.5%
GPT-5.6 Sol23.7%32.6%8.561.27%23.7%
Grok 4.331.6%40.6%9.564.10%52.6%
Kimi K2.623.3%37.1%9.042.42%39.8%
DeepSeek V4 Pro28.6%47.5%8.232.07%30.1%
Mistral Large 333.1%35.3%11.775.42%79.7%
Lower is better. Best result in each column in bold. Evaluation covers 133–135 cases per model.

There is no single “most reliable” model

The results illustrate why evaluating professional AI through a single accuracy number is problematic.

The lowest measured rate of misleading divergence is 17.0%. The lowest rate of false representations is just 0.73% of evaluated statements. The lowest measured incomplete score is 24.3%. But these results belong to different models.

No model performs best across every dimension.

More importantly, apparently small error rates can look very different when we examine the complete work product.

GPT-6 Sol, for example, produces false representations in only 0.84% of evaluated statements. Yet 12.6% of its complete case analyses contain at least one false statement.

GPT-6 Astra reduces the statement-level rate further, to 0.73%, while at least one false statement is still detected in 21.5% of case analyses.

Claude Opus 5.5 produces the lowest measured incomplete score, but 67.2% of its case analyses contain at least one false representation.

These are very different reliability profiles. And that distinction matters because professionals do not rely on average statements.

They rely on the work product in front of them.

Better models make verification more important, not less

As frontier models improve, many obvious failures are becoming less frequent. But the failures that remain can be considerably harder to recognise.

A fabricated citation can be checked. An invented fact can be challenged. An omitted qualification that subtly changes the interpretation of a judgment is much harder to notice, particularly when everything the model did say is correct.

This becomes more important as AI moves from answering individual questions to conducting research, analysing documents and executing longer professional workflows.

From research to our platform

Today, we are bringing this research into our platform for expert work. Our goal is to enable AI-generated professional outputs to be tested against the sources on which they rely, across the four dimensions of failure identified through our research.

We deliberately separate generation from verification. A model generates the work. A separate evaluation layer tests the resulting output against its sources before a professional relies on it.

The objective is not to automate professional judgement or to assume that every legal question has one correct answer. It is to distinguish legitimate differences in professional judgement from failures that violate professional standards and can be systematically identified: misleading, false, incomplete and deficient reasoning.

This is central to our mission:

Making AI-generated professional work reliable and verifiable.

The next era of expert AI won’t belong to the models that generate answers the fastest or sound the most persuasive. It will belong to the systems that give human professionals the rigorous, source-anchored proof required to trust the work product.

We are building the trust layer between raw frontier AI capabilities and professional accountability, giving expert teams the infrastructure to verify every brief, memo, and analysis against its underlying authority before a human expert signs off.

As we continue expanding our evaluation layer, we are partnering with a select group of enterprise professional services firms to test and refine verification across use cases where precision is required.

Request private preview access

We’re working with a small group of professional services firms on verification for their own workflows.

Join the waitlist

Technical report: coming soon.