Census surname field guide

What changed across the 1990, 2000, 2010, and 2020 surname files

The four files sit three decades apart, but they are not four editions of one unchanged table. Their populations, publication rules, editing, fields, and privacy treatments differ. A useful comparison starts by preserving those differences.

1990 is a sample-frequency product

The 1990 file is based on valid surname records in the Post-Enumeration Survey Search Area sample. It reports edited surname spelling, sample frequency percentage, cumulative percentage, and rank. It does not provide the national count that later files publish.

Because the sample design and demographic oversampling affect interpretation, this site displays 1990 as its own measure family. The sample frequency can be inspected, but it is not converted into a later-style national count or placed on a count chart with 2000 through 2020.

2000 establishes the later table shape

The Census 2000 release publishes surnames with a frequency of at least 100 and supplies rank, national count, proportion per 100,000, cumulative proportion, and aggregate race and Hispanic-origin percentages. Its methodology also documents name editing and suppression of demographic cells based on one through four records.

For the core magnitude measures, 2000 is the first vintage on this site with a source-provided national count. It can be placed beside 2010 where the field meanings match, while its processing notes remain attached to the observation.

2010 keeps familiar fields but changes review

The 2010 table uses the same main fields and the same minimum frequency as the 2000 public product, which makes count and rate comparisons possible. Its methodology nevertheless describes differences in editing and a much broader manual scan of candidate names. The published table also contains an “ALL OTHER NAMES” aggregate that is not an individual surname and is excluded here.

A matched field name is necessary for comparison, but it is not a reason to hide these production differences. The site treats 2000 and 2010 counts as comparable published measures, not as a controlled experiment that isolates why a value changed.

2020 uses a protected comparison cohort

The 2020 last-name product limits publication to edited names that met the underlying frequency condition in both 2010 and 2020. It adds disclosure-protection noise to six race and Hispanic-origin component counts. Rank, total, rate, and cumulative proportion are calculated from those protected components.

That construction matters twice. First, the eligible set is narrower than a one-year threshold would be. Second, the displayed total is a derived, noise-affected measure rather than the same kind of source count used in 2000 and 2010. The application exposes both facts wherever it displays a 2020 value.

What a defensible comparison sounds like

A careful comparison names the measure and the vintages: the published count was higher in one later file than another, or the published rate differed between two eligible rows. It also notes when 2020 protection or cohort rules could matter. The observation describes the files; it does not by itself explain migration, marriage, naming practices, data processing, or population change.

The safest conclusion may be a limit. If one endpoint is unpublished, if the measures differ, or if a proposed explanation needs evidence outside the surname tables, stop there. The files are useful precisely because their boundaries can be stated clearly.

  • Compare 2000, 2010, and 2020 counts only where all selected rows are published.
  • Keep 1990 sample frequency outside the later count series.
  • Name the 2020 cohort and disclosure treatment when interpreting that vintage.
  • Treat any explanation for change as a separate research question requiring separate sources.

Research basis

Official sources used for this guide

The interpretation above is original editorial work. Factual claims about the Census products are grounded in these first-party data and methodology artifacts.