Census surname field guide

How to read surname rank, count, and frequency

A surname row can contain several numbers that look interchangeable but answer different questions. The safest reading order is publication status first, method second, and value third. That order prevents a rank from becoming a population estimate or a missing row from becoming a zero.

Start with the published universe

Every rank and rate belongs to a particular source file. The 1990 product describes an edited Post-Enumeration Survey Search Area sample. The 2000 and 2010 files publish surnames meeting a minimum national frequency. The 2020 workbook uses a cohort defined by underlying 2010 and 2020 frequency and then applies disclosure protection. A value is meaningful only inside the universe that produced it.

This is why Surname Time Machine keeps a method note beside every vintage. Before asking whether a surname moved, first ask whether the two files included names under the same rule and measured the same population. If the answer is no, the difference is descriptive, not a clean estimate of one cause.

Rank is an ordering, not a population share

National rank places a published surname relative to the other surnames in that vintage's published universe. A smaller rank number means the name appears earlier in that ordering. It does not say what fraction of the country had the name, and it does not establish that the underlying population rose or fell between two files.

Rank can change even when a surname's own count changes little, because every other published surname also affects the ordering. The eligible set can change as well. Use rank to answer a narrow question—where did this name sit in this file?—and use count or rate for questions about the published magnitude.

Count is available only where the source supports it

The 2000 and 2010 tables provide a national count for each published surname row. The 1990 file does not provide a national surname count, so this site does not estimate one. It reports the sample frequency and rank that the source actually supplies.

The 2020 workbook is different again. Its displayed total is calculated from published race and Hispanic-origin component counts that have been affected by disclosure-protection noise. Surname Time Machine labels that total as derived from published components rather than presenting it as if it had the same construction as the earlier source counts.

Rate provides scale, but not perfect comparability

The later files report a proportion per 100,000. That rate is useful because it puts a surname count in relation to the relevant population total. It is often more informative than count alone when the national population differs between vintages.

A rate does not erase method changes. The underlying universe, editing, eligibility, and disclosure treatment still matter. The site therefore compares later rates only as published vintage-specific measures and keeps 1990's sample frequency percentage separate. Converting one into the other would imply a level of harmonization the sources do not support.

A practical reading sequence

When opening a surname profile, use the same sequence each time. Confirm whether the name was published for the vintage. Read the method note. Choose one matched measure. Compare only the vintages where that measure exists. Then read the sources before drawing a conclusion.

This sequence produces modest but defensible statements: a name had a particular published rank in one file, a published count in another, or a gap under a stated publication rule. It does not support a claim about an individual family, a cause of change, or an unobserved value.

  • Publication status answers whether the file supplies a row.
  • Rank answers position within that file's published ordering.
  • Count answers published magnitude only when a supported count exists.
  • Rate answers published magnitude relative to the vintage population measure.

Research basis

Official sources used for this guide

The interpretation above is original editorial work. Factual claims about the Census products are grounded in these first-party data and methodology artifacts.