Before 2011 the digits were not random. A state and a birth date fix the five-digit prefix — of the time. The sixth digit is real signal on top of that, the seventh is faint, and after it there is nothing left to find: the last two digits of an SSN are as good as coin flips.
574-03-4███
A guess counts here only if every digit up to that point is right, so naming the sixth digit over a wrong prefix earns nothing — that is the coincidence a per-digit score would reward. Measured that way the accuracy falls by roughly a decade per digit, and so does a blind guess: the gap between the two lines is what the birth date actually buys, and it stops widening after the sixth digit.
| Digits | Correct | As the data sits | Blind guess |
|---|
“Correct” weights birth years equally; “as the data sits” counts every record in the source database once.
The clearer test is what each further digit adds. Of the guesses that already have the prefix right, how many survive the next digit? A digit carrying no information survives 10% of the time. The sixth survives —, the seventh —, and the eighth — — which is chance.
Two serials share a leading digit only when they fall in the same block of 1,000, so given a correct prefix the sixth digit is not a gradual reward for a close fit. It is decided by whether the estimate lands within a few hundred numbers of the truth — and a miss of 1,000 to 9,999 cannot produce it at all.
The averages hide how uneven this is. Two things drive it and they point the same way: how many area numbers a state was ever assigned — one for Wyoming, fifty-four for California — and how recently someone was born, because enumeration at birth tied issuance ever more tightly to the birth date.
Where the two meet, the number stops being private. For the eighteen jurisdictions ever given five or fewer area numbers — the small states plus DC and the territories — births from 2005 on (— records), the prefix is right — of the time, the first six digits —, and the first seven —. Given a correct prefix the sixth digit follows — of the time.
— is the strongest single case at — on six digits. Those per-state cells rest on a few hundred records each, so the pooled figure is the one to quote and the ordering between them is approximate.
It comes down to how many prefixes the forty nearest birthdays are spread across. Where those records sit in two or three prefixes, the fit resolves a position inside one of them. Where they sit in twenty, several prefixes were being filled at once and even the prefix is a coin toss.
— Death Master File records with a state, a birth date between 1988 and 2011, and a full nine-digit number. Ten-fold cross-validation: every record is predicted from the nine folds that exclude it, using the nearest forty training birthdays within 180 days, weighted by date distance, over sequence positions where serials count up from 0001 inside each prefix. No record contributes to its own prediction.
Rates are reported two ways because the source is a death file, not a birth cohort: its 1988 births have had twenty-three more years to appear in it than its 2011 births, and early births are the hardest years, so counting every record once understates the rate for a living person. The headline figures weight birth years equally, which is close to a birth-cohort target since US annual births moved less than 10% across this span. The state mix is still the database’s rather than a census’s.
This measures interpolation within one sample of people who have died, not accuracy
against the living population. Full numbers, including per-state, per-year and per-cell
breakdowns, live in the repository’s reports/serial-digit.json.