Section
Litigation Outcome Data
Every number anyone quotes about how cases turn out came from a record that was created for a different purpose. Docketing systems exist to run cases. That they can be counted at all is a side effect, and it shows.
Administrative byproduct, not measurement
A clerk entering a docket text is recording that a thing happened so the case can proceed. Nobody in that workflow is producing a research variable. The consequences run all the way through any analysis built on top: categories exist because they were administratively useful, they change when court practice changes, and they carry no obligation to stay comparable across districts or across decades.
This is not an argument against using the records. It is an argument for reading their documentation before reading their contents, which is close to the whole discipline.
Three different objects get called the same thing
Confusion in this area usually traces back to treating three distinct artefacts as interchangeable.
The docket is the event log for a single case. It is rich, it is narrative, and it is close to the ground. It also has no schema worth the name: the same event is described differently by different courts and sometimes by different people in the same court.
The case-level record is what you get when an administrative body reduces that log to fields. Filing date, nature of suit, disposition, and so on. Anyone who wants to count rather than to read should begin here, for one reason: the reduction has already been performed deliberately and written down, instead of being improvised by whoever happens to be parsing. A documented compression can be argued with. An undocumented one cannot.
The published statistic is a table produced from those records by the Administrative Office of the U.S. Courts. It is authoritative for what it reports and it has already made every aggregation decision on your behalf, which is fine until your question differs slightly from the one the table answers.
Most bad analysis in this space involves quietly substituting one of these for another. A docket search is not a case count. A case count is not an outcome rate.
The selection problem does not go away
The deepest issue is not coding quality. It is that the cases visible in any court dataset are the ones that reached a court and stayed there long enough to leave a trace. Disputes resolved before filing are absent by construction. So, largely, are the terms of settlements, which are frequently the actual outcome and are frequently confidential.
A dataset of filed cases can tell you a great deal about filed cases. Reasoning from it to a claim about disputes in general requires an argument about why the filtering does not matter, and that argument almost never appears alongside the number.
When two sources disagree
Sooner or later a figure taken from one of these artefacts will contradict a figure taken from another, and the instinct to work out which is right is the wrong instinct. In this domain the answer is almost always that both are correct and they are counting different things.
Four differences account for most of it. The population may differ: filed cases, terminated cases, and pending cases are three separate denominators, and a table reporting one is routinely compared against a query returning another. The period may differ, because a judicial year is not a calendar year. The unit may differ, since one dispute can be several dockets after a transfer or a consolidation, and a count of dockets is not a count of matters. And the category definition may have changed inside the window being compared, which nothing in the output will announce.
The productive order is to reconcile definitions before touching data. Read what each source says it is counting, work out which of the four differences applies, and only then decide whether there is a discrepancy left to explain. In the large majority of cases there is not, and the exercise has taught you what your own number actually means. Where something does survive that process, it is worth reporting to whoever publishes the file rather than quietly working around.
What is here
The first article reads a federal civil docket as a data source: what each entry type actually commits to, which fields are reliable, and which questions the docket alone cannot settle no matter how carefully it is parsed. The second takes disposition coding seriously, working through what the Integrated Data Base records when a case ends and where that compression misleads a reader who has not seen the code book.
Both assume you have looked at a docket before. Neither assumes any statistics background. Where the analysis in these pieces meets the tooling question, the explainability section covers the other side of it, and the rest of this publication shows how the sections relate.