Census surname field guide

How to read the 2020 disclosure-protected surname table

The 2020 surname workbook looks familiar if you have used the 2010 table, but its rows and numbers are produced under a different publication design. Eligibility depends on both censuses, published components receive disclosure protection, and several familiar-looking fields are calculated from those components. Reading the workbook responsibly starts with that construction rather than treating it as a simple new count list.

The workbook is a two-vintage cohort

The public last-name product is limited to edited surnames that met the underlying frequency condition in both 2010 and 2020. A row is therefore evidence that the name qualified for this comparison cohort, not evidence that the workbook contains every surname that independently met a one-year threshold in 2020. The cohort rule shapes both what appears and what remains absent.

This distinction matters when a name cannot be found. The missing row does not prove that nobody used the surname in 2020, and it does not identify which side of the two-vintage condition prevented publication. A careful result reports the name as unpublished or ineligible for this public product and stops before assigning a numerical value.

Protection is applied to published components

The workbook publishes six race and Hispanic-origin component counts for each surname row. The Census methodology describes noise infusion and additional processing applied to protect those components. Surname Time Machine preserves a noise-affected status on the 2020 observation so the protection remains visible wherever the result is displayed.

Disclosure protection does not license a generic claim that every displayed value differs from an unobserved value by one fixed amount. The method must support any statement about uncertainty. This site therefore states the construction and avoids attaching a universal plus-or-minus range to every surname total.

The displayed total is derived

The public total is calculated from the six published protected components rather than supplied as the same kind of untouched count field used in the earlier tables. Rank, proportion per 100,000, and cumulative proportion are then based on the protected output. The result is still a published Census measure, but its count semantics differ from 2000 and 2010.

Protected components can produce totals below the unprotected eligibility threshold after processing. That does not mean the surname failed the rule or that the workbook contradicts itself. Eligibility and public output occur at different stages. A citation should describe the published total without using it to reverse-engineer the confidential value that preceded protection.

Comparison with 2010 requires matched questions

The 2020 cohort is designed around names present under the underlying condition in both vintages, which makes a 2010 comparison possible for the included rows. Even so, the two displayed count fields do not have identical construction. The 2010 table supplies a published source count, while the 2020 result is derived from protected components.

A responsible comparison can report the two published values and their documented semantics. It cannot automatically explain the difference as migration, births, marriage, spelling changes, or cultural change. Those are separate causal questions requiring evidence beyond the surname tables. Rank movement also remains distinct from count or rate movement because the surrounding published ordering can change.

Use a five-step reading check

First, confirm that the exact written surname has a published row. Second, identify the displayed field rather than referring vaguely to “frequency.” Third, note whether the value is a component, derived total, rank, or proportion. Fourth, carry the disclosure-protection caveat into any comparison. Fifth, check the 2010 row separately instead of assuming a missing or matching value.

This sequence keeps the useful information intact while avoiding claims the workbook cannot support. You can say what the protected public product reports, where the name ranks within that product, and how its published measures compare under named conditions. You cannot use the row to identify an individual, infer ancestry, or recover an exact unprotected count.

  • Treat eligibility and published protected output as separate stages.
  • Label the total as derived from protected components.
  • Compare 2010 and 2020 only with the construction difference visible.
  • Do not infer a person’s identity or an unprotected value from an aggregate surname row.

Research basis

Official sources used for this guide

The interpretation above is original editorial work. Factual claims about the Census products are grounded in these first-party data and methodology artifacts.