Legal-Pythia

Legal technology | Litigation data | Public datasets | Primary sources, dated | Not a law firm

How Case Outcomes Are Coded in Public Data

Litigation Outcome Data 2026-06-11 2 primary sources Sources checked 2026-08-28

When a federal civil case ends, the whole of it gets reduced to a handful of coded fields. The Federal Judicial Center publishes those fields for civil and criminal cases in its Integrated Data Base, along with the documentation describing what each code means.

That documentation is the article. Everything below is an argument for reading the code book before the data, and a description of what goes wrong when people do not.

Compression is the design, not a defect

A case that ran four years, generated three hundred docket entries, and settled on terms nobody outside the room knows becomes a filing date, a termination date, a nature-of-suit category, a disposition code, and a small number of related fields.

This is not carelessness. The purpose of the record is to let the judiciary describe and manage its own workload, and for that purpose the compression is appropriate. The trouble starts when a field built to support caseload management is read as a finding about who was right.

Nature of suit is a filing-time guess

The nature-of-suit category is what a case was labelled as at the start. It is chosen when the matter is initiated, from a fixed list, and it does not follow the case as claims are amended, dropped, or added.

Two consequences follow and both are routinely ignored. A case that began as one kind of dispute and became another retains its original label, so any analysis by category is analysing initial framing rather than substance. And because the list is fixed, disputes that do not fit it neatly land in whichever bucket is closest, which means the residual and general categories are heterogeneous in ways their names do not advertise.

Anyone comparing categories across a long period has an additional problem: the list itself has been revised. A series that looks continuous may not be.

Disposition is a procedural fact

The disposition field records how the case left the court. The available values distinguish things like transfer, remand, dismissal, and judgment, and they are accurate about the procedural route.

They are close to silent about the merits. Consider what a dismissal covers. It includes a case thrown out because the claim failed as a matter of law. It also includes a case dismissed because the parties settled and asked the court to close the file, which is frequently a plaintiff outcome and is coded identically to the first. Reading dismissal as a defence win conflates the two, and the direction of that error is not small.

Judgment codes have the mirror problem. A judgment entered after the other side stopped participating is procedurally a judgment and substantively a case that was never contested. The field will not distinguish that from a judgment after trial, because the field was not built to.

Amounts, where they exist at all

A monetary field exists and it is much less useful than it looks. It reflects what was recorded in the closing documents, which is not the same as what changed hands, and it is systematically absent for confidential settlements. Because those settlements are not distributed randomly across case types or sizes, the available amounts are a biased sample and the bias runs in a direction that depends on the category.

A defensible use of the field is describing the recorded amounts and saying so explicitly. An indefensible one is presenting their distribution as the distribution of case values.

How to work with it anyway

The records are genuinely valuable and the way to use them is to keep the questions inside what the coding supports.

  • Read the code book for the specific years in scope, not a general description of the database.
  • Ask what an unusual value means before excluding it. Missing and not applicable are different, and both are informative.
  • Check whether a category's definition changed inside the period being compared.
  • State the denominator every time, because filed, terminated, and pending are three different populations.
  • Where the published tables from the Administrative Office of the U.S. Courts already answer the question, use them and cite them rather than rebuilding the aggregation.

That last point saves more errors than any other. The published tables have made their aggregation decisions deliberately and documented them. A privately rebuilt version of the same statistic that disagrees is far more likely to be wrong than to be a discovery, and the first move on finding a discrepancy is to work out which aggregation choice produced it.

The preceding article covers the event-level records these codes are derived from. For the method guides applying the same reading to datasets outside the courts, the rest of this publication is the route through.


Primary sources

  1. Integrated Data Base Federal Judicial Center
  2. Caseload statistics data tables Administrative Office of the U.S. Courts