Census surname field guide
What changed across the 1990, 2000, 2010, and 2020 surname files
The four files sit three decades apart, but they are not four editions of one unchanged table. Their populations, publication rules, editing, fields, and privacy treatments differ. A useful comparison starts by preserving those differences.
1990 is a sample-frequency product
The 1990 file is based on valid surname records in the Post-Enumeration Survey Search Area sample. It reports edited surname spelling, sample frequency percentage, cumulative percentage, and rank. It does not provide the national count that later files publish.
Because the sample design and demographic oversampling affect interpretation, this site displays 1990 as its own measure family. The sample frequency can be inspected, but it is not converted into a later-style national count or placed on a count chart with 2000 through 2020.
2000 establishes the later table shape
The Census 2000 release publishes surnames with a frequency of at least 100 and supplies rank, national count, proportion per 100,000, cumulative proportion, and aggregate race and Hispanic-origin percentages. Its methodology also documents name editing and suppression of demographic cells based on one through four records.
For the core magnitude measures, 2000 is the first vintage on this site with a source-provided national count. It can be placed beside 2010 where the field meanings match, while its processing notes remain attached to the observation.
2010 keeps familiar fields but changes review
The 2010 table uses the same main fields and the same minimum frequency as the 2000 public product, which makes count and rate comparisons possible. Its methodology nevertheless describes differences in editing and a much broader manual scan of candidate names. The published table also contains an “ALL OTHER NAMES” aggregate that is not an individual surname and is excluded here.
A matched field name is necessary for comparison, but it is not a reason to hide these production differences. The site treats 2000 and 2010 counts as comparable published measures, not as a controlled experiment that isolates why a value changed.
2020 uses a protected comparison cohort
The 2020 last-name product limits publication to edited names that met the underlying frequency condition in both 2010 and 2020. It adds disclosure-protection noise to six race and Hispanic-origin component counts. Rank, total, rate, and cumulative proportion are calculated from those protected components.
That construction matters twice. First, the eligible set is narrower than a one-year threshold would be. Second, the displayed total is a derived, noise-affected measure rather than the same kind of source count used in 2000 and 2010. The application exposes both facts wherever it displays a 2020 value.
What a defensible comparison sounds like
A careful comparison names the measure and the vintages: the published count was higher in one later file than another, or the published rate differed between two eligible rows. It also notes when 2020 protection or cohort rules could matter. The observation describes the files; it does not by itself explain migration, marriage, naming practices, data processing, or population change.
The safest conclusion may be a limit. If one endpoint is unpublished, if the measures differ, or if a proposed explanation needs evidence outside the surname tables, stop there. The files are useful precisely because their boundaries can be stated clearly.
- Compare 2000, 2010, and 2020 counts only where all selected rows are published.
- Keep 1990 sample frequency outside the later count series.
- Name the 2020 cohort and disclosure treatment when interpreting that vintage.
- Treat any explanation for change as a separate research question requiring separate sources.
Research basis
Official sources used for this guide
The interpretation above is original editorial work. Factual claims about the Census products are grounded in these first-party data and methodology artifacts.
- 1990 Census Frequently Occurring Surnames (dist.all.last)U.S. Census Bureau. Sample-based file; it supplies frequency percent, cumulative frequency percent, and rank, but no national count.Read the methodology
- Documentation and Methodology for Frequently Occurring Names in the U.S. - 1990U.S. Census Bureau. Documents the 1990 Post-Enumeration Survey Search Area sample, editing, missingness, coverage, and limitations.
- Demographic Aspects of Surnames from Census 2000U.S. Census Bureau. Documents the Census 2000 surname universe, the 100-occurrence publication threshold, editing, and cell suppression for values from 1 through 4.
- Frequently Occurring Surnames in the 2010 CensusU.S. Census Bureau. Documents 2010 edits, comparison limits, and the 162,253-name public table covering names with frequency at least 100.
- Last Name Data From the 2020 CensusU.S. Census Bureau. Documents the 2010-and-2020 cohort rule, noise infusion into race and Hispanic-origin counts, derived totals, optimization against negatives, and the possibility of published totals below 100.