Census surname field guide
Why the 1990 Census surname file must stand apart
The 1990 surname file is valuable historical evidence, but it is not an earlier copy of the later national surname tables. It comes from a sample, uses its own editing and coverage rules, and publishes frequency percentage, cumulative frequency percentage, and rank without a national surname count. Keeping it in a separate analytical frame is the key to using it accurately.
The source begins with a sample
The 1990 product was built from records associated with the Post-Enumeration Survey Search Area sample described in its methodology. That sampling context defines what the file can represent. It should not be described as a complete enumeration of every surname used by the national population, and its rows should not be assumed to share the later tables’ publication universe.
Sampling also changes how absence should be read. A surname missing from the file may have been outside the observed or released material. The absence does not support a claim that the name was unused in 1990, had a national count of zero, first appeared later, or disappeared between two Census products.
Three fields are published—and count is not one of them
Each surname row supplies a frequency percentage, a cumulative frequency percentage, and a rank. Frequency percentage describes the surname’s share within the source’s documented sample framework. Cumulative frequency describes the running total through the ranked file. Rank identifies the row’s ordering within that published material.
The file does not publish a national surname count. Multiplying the sample percentage by a later population estimate would introduce assumptions about the sample, denominator, weighting, editing, and target population that the row does not provide. Surname Time Machine stores the 1990 count as unavailable and never draws a count bar for that vintage.
Rank remains local to the file
A 1990 rank answers where the surname sits in the ordering of this particular file. The rank can be compared descriptively with a later published rank only when the change in universe and method is stated. A smaller or larger number does not by itself measure population growth, decline, cultural prominence, or the experience of an individual family.
Rank is affected by every other row in the ordering as well as by the surname’s own observed frequency. Later products use different source populations and publication rules, so a rank difference combines multiple moving parts. The defensible statement names both ranks and both files rather than translating the difference into an unexplained trend label.
Later count charts should begin in 2000
The 2000 and 2010 public surname tables include published national counts for qualifying rows. The 2020 product supplies a protected, derived total. Those values have their own differences, but they at least answer count-like questions under documented later methods. The 1990 sample frequency cannot be inserted as a count endpoint without manufacturing a value.
A clear presentation therefore gives 1990 its own sample-frequency panel and begins the count comparison with 2000. The visual gap is intentional. A smooth four-point line would look simpler, but it would imply a continuous measure that the sources do not contain. Accuracy takes priority over visual symmetry.
Ask questions the file can answer
The file can support a statement that an exact surname row has a particular published sample frequency, cumulative frequency, or rank. It can also support research into the structure and coverage of the publication itself. It cannot establish a national count, a personal family history, a state distribution, or a cause for a difference observed in another decade.
When citing 1990, name the sample-based source and the field. When comparing, place the later value beside it rather than converting one into the other. When the row is absent, describe the publication gap. These modest rules preserve the historical value of the file without asking it to answer questions it was not designed to answer.
- Use only frequency percentage, cumulative frequency percentage, and rank from the 1990 row.
- Keep national count explicitly unavailable for every 1990 surname.
- Describe rank within the file rather than as population growth or decline.
- Show a publication gap instead of connecting or estimating a missing observation.
Research basis
Official sources used for this guide
The interpretation above is original editorial work. Factual claims about the Census products are grounded in these first-party data and methodology artifacts.
- 1990 Census Frequently Occurring Surnames (dist.all.last)U.S. Census Bureau. Sample-based file; it supplies frequency percent, cumulative frequency percent, and rank, but no national count.Read the methodology
- Documentation and Methodology for Frequently Occurring Names in the U.S. - 1990U.S. Census Bureau. Documents the 1990 Post-Enumeration Survey Search Area sample, editing, missingness, coverage, and limitations.
- Demographic Aspects of Surnames from Census 2000U.S. Census Bureau. Documents the Census 2000 surname universe, the 100-occurrence publication threshold, editing, and cell suppression for values from 1 through 4.
- Frequently Occurring Surnames in the 2010 CensusU.S. Census Bureau. Documents 2010 edits, comparison limits, and the 162,253-name public table covering names with frequency at least 100.