What the RECAP Archive's Coverage Skews Toward
There is a free, searchable collection of millions of federal court filings, and it is one of the most useful things on the legal internet. It is also not a sample of federal litigation, and the difference matters the moment anyone counts something in it.
How the collection came to exist
The archive describes its own construction without any spin, which is more than most data sources manage:
The RECAP Archive is a searchable collection of millions of PACER documents and dockets that were gathered using our RECAP Extensions for Firefox, Chrome and Safari.
Free Law Project, About CourtListener's RECAP Archive
A browser extension. Somebody paid for a document on the federal courts' own access system, the extension noticed, and a copy went into the archive. Multiply by thousands of contributors over years and you get millions of documents.
That is a genuinely clever solution to a real problem, and the project is explicit that the problem it addresses is the cost of access. It also determines the shape of what exists.
The selection mechanism, stated plainly
A document is in the archive if someone with the extension installed chose to pay for it. So inclusion is a function of three things stacked: whether anyone wanted the document, whether that person happened to be running the extension, and whether the matter was interesting enough to be worth the fee to somebody.
Each of those correlates with the same underlying property. Contested matters attract more lookups than uncontested ones. Large cases attract more than small ones. Newsworthy litigation attracts far more than routine litigation. Cases involving parties that journalists, academics, and researchers follow attract more than cases involving parties nobody is watching, and those are precisely the people most likely to have the extension installed.
The practical result is that the typical case in the archive is more contested, longer, and larger than the typical case in the federal courts. Any statistic computed over the archive and described as a statistic about federal litigation inherits that skew in full.
One systematic component, which cuts the other way
The picture is not purely voluntary, and this is the part most descriptions of the archive miss. It also holds every filing that the federal system makes available at no charge.
That slice is systematic rather than contributed, and it is a specific kind of document rather than a random one. So the archive is a union of two differently shaped collections: a voluntary one skewed toward attention, and a free-of-charge one skewed toward a particular document type. A researcher who models the coverage as one mechanism will be wrong about both halves.
What it is excellent for
None of this makes the archive less valuable. It makes it valuable for different questions than the ones people reach for.
- Finding a specific document in a specific case, which is what it is built for.
- Reading the actual text of filings, including scanned material the project converts to text.
- Studying a defined population you have enumerated independently, then pulling documents for it.
- Any question where you supply the sampling frame and use the archive only as a document source.
That last one is the pattern worth adopting. Decide which cases you care about from a source built to be complete, then use the archive to read them. The failure mode is the reverse: searching the archive, taking what comes back, and treating the result as a population.
Pairing it with something complete
For counting, the Federal Judicial Center's Integrated Database is the better instrument, because it comes from what the courts report administratively rather than from what anyone chose to download. It has its own seams, covered in a companion article, and they are seams of definition rather than of coverage.
The two together are strong. Enumerate from the administrative data, read from the archive, and never compute a rate whose numerator and denominator come from different ones. Half the citation-worthy findings in empirical legal work come from exactly that pairing, and half the embarrassing ones come from using either alone for the job of the other.
A note on quoting coverage numbers
The archive is described as holding millions of documents and dockets, and that is a floor that grows as people contribute. It is not a proportion of federal filings, and a number of documents is not a coverage rate. Anyone wanting a coverage rate has to name a denominator from somewhere else, and then defend the assumption that the two sources define a case the same way, which is the harder half of the exercise.
Say what you have. A count of documents in a contributed archive is a real and useful thing to report, and it is not a claim about the courts.
The neighbouring section covers the tooling questions that arise once you have the documents. For the full scope, the full index of research notes is the way through.
Primary sources
- The RECAP Archive Free Law Project
- CourtListener Free Law Project
- Integrated Data Base Federal Judicial Center