Evaluating Accuracy Claims in Legal AI Tools
A single percentage is the most efficient way to end a conversation that should have continued. It arrives looking like a measurement and it is a summary of a table nobody has shown you.
What follows is the set of things that table would have to contain. None of it is specialised. All of it is routinely absent.
The denominator decides the number
Accuracy is correct decisions over total decisions, so the composition of the total is doing most of the work. In document classification the classes are almost never balanced. If four per cent of a corpus is genuinely responsive, a system that marks every document not responsive scores ninety-six on accuracy while being worthless.
That example sounds too crude to matter and it is the shape of a large fraction of real results. Any figure quoted without the base rate of the thing being detected is uninterpretable rather than merely incomplete, and the base rate is the first thing to ask for.
Your two errors do not cost the same
Once the classes are unbalanced, one aggregate number cannot carry the information, because the two ways of being wrong land in different places.
Missing a responsive document is a disclosure failure. It surfaces late, if it surfaces at all, and it surfaces in front of the people best placed to make it expensive. Flagging an irrelevant document is a review cost. It surfaces immediately and it is measured in hours.
These are so asymmetric that the sensible operating point is usually nowhere near the one that maximises accuracy. A tool tuned to catch nearly everything, at the price of a large pile of false positives, may be exactly right for privilege screening and exactly wrong for a first-pass cull. Neither configuration is better in the abstract, and a supplier quoting one number has picked an operating point without telling you which.
The useful request is not for a better number. It is for the trade-off curve, and for the supplier to say which point on it their headline figure describes.
Where the test set came from
Every performance claim is a claim about a specific set of documents, and its transferability depends entirely on how that set relates to yours.
Three questions get most of the way there. Was the evaluation set held out completely, or did any part of it inform tuning? Was it drawn from the same matter as the training data, meaning the same custodians, the same period, the same house vocabulary? And who assigned the ground truth labels, under what instructions, with what measured agreement between them?
The last one is regularly the weakest link and almost never discussed. Ground truth in document review is human judgement, and human reviewers disagree with each other at rates that would embarrass most models. A system reported as matching human labels ninety-something per cent of the time, evaluated against labels two humans agreed on far less often than that, is being scored against a ruler that moves.
The knowledge-limits question, again
Every figure discussed so far describes behaviour on inputs resembling the test set. None of them says anything about behaviour on inputs that do not, which is the fourth NIST principle and the one a percentage structurally cannot address.
This matters more in litigation than in most deployments because the input distribution is not chosen. A collection arrives as it arrives. It contains the file formats it contains, in the languages the custodians used, covering a period nobody selected for convenience. Performance measured on a tidy benchmark and performance on that are different quantities, and only one of them is on the slide.
What the regulatory picture adds
Two external frames are worth knowing, less because they settle anything than because they indicate where the questions are heading.
The European Union's Regulation 2024/1689 on artificial intelligence organises obligations by the risk attached to a use rather than by the technique, which puts the burden of classification on the deployer's purpose. Whatever a given system's status under it, the structural point carries: the same model can sit in a low-stakes workflow and a high-stakes one and attract different duties in each. Buying decisions made without reference to the intended use are answering the wrong question.
Closer to the ground, the Federal Rules of Evidence set the conditions under which expert opinion is admitted. Where model output feeds anything approaching testimony, the framing shifts from performance to method: whether the principles are reliable and whether they were reliably applied to the facts at hand. A percentage is not responsive to either. A documented, reproducible, explainable process is.
A short list to take into the room
- What proportion of the evaluation corpus was actually in the positive class.
- Show the trade-off curve, and say which point the quoted figure describes.
- How was the evaluation set separated from everything used in tuning.
- Who produced the ground truth, and how often did they agree with each other.
- Demonstrate the behaviour on an input the system was not built for.
A supplier who can answer these is not necessarily selling a better system, and is definitely selling a more measurable one. That is the property worth paying for, because it is the only one that survives being asked about afterwards by somebody hostile.
The preceding article in this section sets out the four principles these questions come from. For how the underlying records behave once you start counting them, our other work on court data leads to the rest.
Primary sources
- Four Principles of Explainable Artificial Intelligence (NISTIR 8312) National Institute of Standards and Technology Cached in this repository at research/sources/NIST.IR.8312.pdf
- Regulation (EU) 2024/1689 on artificial intelligence Publications Office of the European Union
- Federal Rules of Evidence Administrative Office of the U.S. Courts