Census surname field guide

Why the 1990 Census surname file must stand apart

The 1990 surname file is valuable historical evidence, but it is not an earlier copy of the later national surname tables. It comes from a sample, uses its own editing and coverage rules, and publishes frequency percentage, cumulative frequency percentage, and rank without a national surname count. Keeping it in a separate analytical frame is the key to using it accurately.

The source begins with a sample

The 1990 product was built from records associated with the Post-Enumeration Survey Search Area sample described in its methodology. That sampling context defines what the file can represent. It should not be described as a complete enumeration of every surname used by the national population, and its rows should not be assumed to share the later tables’ publication universe.

Sampling also changes how absence should be read. A surname missing from the file may have been outside the observed or released material. The absence does not support a claim that the name was unused in 1990, had a national count of zero, first appeared later, or disappeared between two Census products.

Three fields are published—and count is not one of them

Each surname row supplies a frequency percentage, a cumulative frequency percentage, and a rank. Frequency percentage describes the surname’s share within the source’s documented sample framework. Cumulative frequency describes the running total through the ranked file. Rank identifies the row’s ordering within that published material.

The file does not publish a national surname count. Multiplying the sample percentage by a later population estimate would introduce assumptions about the sample, denominator, weighting, editing, and target population that the row does not provide. Surname Time Machine stores the 1990 count as unavailable and never draws a count bar for that vintage.

Rank remains local to the file

A 1990 rank answers where the surname sits in the ordering of this particular file. The rank can be compared descriptively with a later published rank only when the change in universe and method is stated. A smaller or larger number does not by itself measure population growth, decline, cultural prominence, or the experience of an individual family.

Rank is affected by every other row in the ordering as well as by the surname’s own observed frequency. Later products use different source populations and publication rules, so a rank difference combines multiple moving parts. The defensible statement names both ranks and both files rather than translating the difference into an unexplained trend label.

Later count charts should begin in 2000

The 2000 and 2010 public surname tables include published national counts for qualifying rows. The 2020 product supplies a protected, derived total. Those values have their own differences, but they at least answer count-like questions under documented later methods. The 1990 sample frequency cannot be inserted as a count endpoint without manufacturing a value.

A clear presentation therefore gives 1990 its own sample-frequency panel and begins the count comparison with 2000. The visual gap is intentional. A smooth four-point line would look simpler, but it would imply a continuous measure that the sources do not contain. Accuracy takes priority over visual symmetry.

Ask questions the file can answer

The file can support a statement that an exact surname row has a particular published sample frequency, cumulative frequency, or rank. It can also support research into the structure and coverage of the publication itself. It cannot establish a national count, a personal family history, a state distribution, or a cause for a difference observed in another decade.

When citing 1990, name the sample-based source and the field. When comparing, place the later value beside it rather than converting one into the other. When the row is absent, describe the publication gap. These modest rules preserve the historical value of the file without asking it to answer questions it was not designed to answer.

  • Use only frequency percentage, cumulative frequency percentage, and rank from the 1990 row.
  • Keep national count explicitly unavailable for every 1990 surname.
  • Describe rank within the file rather than as population growth or decline.
  • Show a publication gap instead of connecting or estimating a missing observation.

Research basis

Official sources used for this guide

The interpretation above is original editorial work. Factual claims about the Census products are grounded in these first-party data and methodology artifacts.