Reconciling a Texas Fatal Crash Count
Ask how many people died on Texas roads last year and you can get two authoritative answers that do not match. One comes from the state, which is the custodian of its own crash records. The other comes from the federal fatality census. Neither is wrong, and the instinct to determine which one is correct will waste an afternoon.
Four documented differences account for nearly all of the gap. Working through them in order is faster than any amount of staring at the two numbers.
One: the survival window
The federal census counts a death only where it occurred within thirty days of the crash. That is a stated inclusion criterion, not a convention, and it exists so files can close on a schedule.
A state record system is built to hold crash reports rather than to close a fatality count, so a death recorded against a crash need not sit inside the same window. Where the two treat a late death differently, the federal figure is the lower one. By how much is unknowable from either file, because the cases that would measure the difference are exactly the ones the narrower definition declines to record, which is a nice illustration of why you cannot audit a dataset using only itself.
Two: what counts as a road
The federal criterion turns on whether the surface was customarily open to public travel. The Texas one is drawn around the road network the state is responsible for, which is a question about jurisdiction rather than about access.
Those are different tests aimed at the same intuition, written by different bodies for different reasons, and neither was drafted with the other in mind. A crash on a road that is open to anyone but not part of the network the state maintains can satisfy one test and fail the other, and nothing in either file flags that it sat near a boundary.
Three: what triggers a record at all
This is the largest structural difference and it runs in the opposite direction from the first two.
Texas records exist because an officer investigated and filed a report, with the threshold set at injury, death, or apparent property damage of a stated amount. Federal fatality records exist because a death occurred and met the census criteria. For fatal crashes specifically the two mostly converge, because a fatal crash is investigated. For everything below fatal they do not converge at all, since the federal census does not attempt non-fatal crashes.
So the comparison is only meaningful for fatal crashes in the first place. Anyone comparing total crash counts between the two is comparing a fatality census against a general reporting system, which is not a discrepancy but a category error.
Four: when the file was closed
Both systems finalise records over time, and they do not finalise on the same clock. A death reported late, a report amended, or a case reclassified will land in one publication cycle for one system and a different cycle for the other. This means two extracts pulled on the same day can still disagree about a year that closed long ago, and it means the disagreement can shrink on its own without anyone doing anything. Analysts new to this pair usually discover it by rerunning a query and getting a different answer.
Texas additionally holds data on a rolling retention window of ten years plus the current year, which means the same query run at two different times covers different periods. The federal files stay available back to 1975. For any comparison, both the vintage of the extract and the date it was pulled belong in the citation, and the absence of that stamp is the single most common reason two analysts cannot reproduce each other's numbers.
A reconciliation that actually works
- Restrict to fatal crashes. Below that the two systems are not attempting the same task.
- Fix the geography to the same definition, and say which one.
- Apply the thirty-day rule to the state data if you can, or state that you could not and that the state figure is therefore the broader one.
- Record the extract date for both, and the retention window for the state figure.
- Report the residual difference rather than reconciling it away. A small unexplained remainder is an honest result.
Done in that order the gap usually shrinks to something small and explicable. It will not close completely. A write-up claiming it did has either got lucky or stopped reporting something.
Which to use
For a national comparison or a long series, the federal census, because it applies one definition across every state and reaches back five decades. For anything about Texas roads below the fatal threshold, the state system, because the federal census does not contain those crashes. For a specific carrier or vehicle type, neither on its own.
And for a question phrased as how dangerous a road is, both, separately, with the definitions quoted. A single number with no definition attached is the thing this pair of datasets is unusually good at teaching you not to publish.
The two companion guides in this section cover each of these systems on its own terms, and reading them before attempting a reconciliation will save most of the work described above. Everything published here sits in the full index of research notes.
Primary sources
- CRIS public query tool Texas Department of Transportation
- Crash reports and records Texas Department of Transportation
- Fatality Analysis Reporting System National Highway Traffic Safety Administration
- Fatality Analysis Reporting System Analytical User's Manual, 1975-2024 (DOT HS 813794) National Highway Traffic Safety Administration Cached in this repository at research/sources/FARS-813794.pdf