Probabilistic Race and Ethnicity Prediction Using Group-Specific Name Lists
Authors: Kyla Chasalow, Noah Dasanaike, Kosuke Imai
Organizations: PhD Candidate, Department of Statistics, Harvard University. · PhD Candidate, Department of Government, Harvard University. · Edith and Benjamin Geisinger Professor, Department of Government and Department of Statistics, Harvard University. 1737 Cambridge Street, Institute for Quantitative Social Science, Cambridge MA 02138.
Statistically valid estimation of racial and ethnic disparities often requires inferring the probability that an individual belongs to a particular racial or ethnic group given only their name and geographic location. The standard approach, Bayesian Improved Surname Geocoding (BISG), relies on group population frequencies for each name. Although the U.S. Census Bureau provides such information for common names and a limited set of racial categories, comparable data do not exist for many racial and ethnic groups and are rarely available outside the U.S. We propose the list-powered BISG (ℓBISG) method, which can be used to derive calibrated group probabilities from group-specific name lists. These lists may be compiled based on expert knowledge or generated synthetically using large language models (LLMs), and thus may be subject to unknown biases. Representing names as embeddings, we treat list membership as a proxy prediction task and apply a correction based on proximal inference to recover the target group probabilities. We validate the method on U.S. voter files with self-reported race, on the full-count 1900 U.S. Census, and on the Lebanese voter registry. We find that LLM-generated name lists yield accurate and well-calibrated probabilities as well as precise disparity estimates comparable to those obtained using methods that require name-race data. Thus, ℓBISG substantially broadens the applicability of probabilistic race and ethnicity prediction to settings where name-race data are unavailable.
Figures & tables
Target dataset of n observations {Si,Gi}i=1n ,
Geographic prevalence estimates Pr(R=r∣G=g) ,
Name lists Lr(S) for every group r=1,…,dR ,
Pre-trained embedding model E(S) ,
Number of folds J≥2 ,
Number of clusters K≥2 .
Algorithm 1 ℓ BISG
List-score vectors f^Ei and geographies Gi from Step 5 of Algorithm 1 ,
Geographic prior Pr(R=r∣G=g) ,
Candidate set K for the number of clusters,
Number of folds J≥3
Algorithm 2 Choosing the number of clusters K
Figure 1 : Florida surnames plotted by log commonness Pr(S∣R) ( x -axis) against distinctiveness Pr(R∣S) ( y -axis) for the four race groups. Each point is a surname held by at least twenty-five members of the group, sized by how many voters hold it. Colored points are those included in the list of 1,000 surnames generated by an LLM for that group. Labeled points are the most common surnames in each group. The dashed line marks the distinctiveness floor of 0.3 used to build lists in Section 5.1 .
Coverage
Precision
Group
List
In Census
Ground truth
Estimate
Ground truth
Estimate
White
Census
100%
0.28
0.30
0.79
0.74
LLM
100%
0.31
0.32
0.71
0.64
Black
Census
100%
0.33
0.37
0.46
0.46
LLM
100%
0.53
0.57
0.25
0.24
Hispanic
Census
100%
0.50
0.53
0.82
0.97
Table 1 : Comparison of LLM-based and Census-based surname lists on the Florida registrants. The column labeled as “In Census” shows the share of list names that appear in the 2020 Census surname table at all. “Coverage” is the share of the group’s registrants with a surname on the list, and “Precision” the share of the registrants with a list surname who belong to the group, each computed from the self-reported race (ground truth) and estimated from the tract system with the tract-level race shares as the geographic prior and no individual labels (Section 4.3 ).
Figure 2 : Precision and recall for ℓ BISG (LLM lists, red, solid), standard BISG (blue, dashed) and geography alone (grey, dotted), by state (rows) and race (columns). Average precision is reported above each panel.
Figure 3 : Calibration of estimated individual race probabilities based on ℓ BISG (LLM lists, red, solid) and standard BISG (blue, dashed) for Florida (top row) and North Carolina (bottom row). The dashed diagonal marks perfect agreement between estimated probabilities and observed proportions. The area of each circle is proportional to the number of registrants in the bin.
Figure 4 : County-level race-specific estimates of the Democratic share of registrants. Rows give signed bias (top) and root-mean-square error (bottom), both in percentage points and averaged over counties. The left two columns are Florida and the right two North Carolina, each pair showing the surname-only estimators (BISG) and then those that add first names (BIFSG). Red bars are the ℓ BISG/ ℓ BIFSG (LLM lists) estimators, while blue bars are the estimates based on the standard BISG/BIFSG.
Coverage
Precision
List
Records
Truth
Oracle
Coarse
Truth
Oracle
Coarse
Chinese
114,090
0.55
0.65
0.51
0.19
0.21
0.17
Japanese
87,278
0.26
0.27
0.24
0.71
0.72
0.67
White
66,834,064
0.33
0.32
0.32
0.83
0.80
0.80
Black
8,717,147
0.59
0.74
0.74
0.20
0.25
0.25
American Indian
34,207
0.31
0.22
0.21
0.00
0.00
0.00
Table 2 : Quality of the five LLM lists on the full-count 1900 census. “Records” counts the true number of people enumerated in each group. “Coverage” is the share of those records with a surname on the group’s list, and “Precision” the share of all people with a surname on the list who belong to the group. “Truth” is computed from the enumerated race and ethnicity. “Oracle” is estimated by the method of Section 4.3 with the true county prevalence of every group. “Coarse” uses only the coarse geography: for the Chinese and Japanese lists, coverage is obtained from the county-level list hit rates and the county Asian share by the method of Section 3 , and precision is the estimated coverage times the recovered number of subgroup members, as a share of the people holding a list surname; for the other lists, whose county prevalence is known, it equals the method of Section 4.3 .
Figure 5 : Precision and recall (top) and calibration (bottom) over all records of the census for coarse prior ℓ BISG, which recovers the subgroup geography from the coarse geography and the lists by the method of Section 3.3 (red, solid); oracle ℓ BISG, which uses the true subgroup geography (blue, dashed); and, in the top panel, recovered geography, the county subgroup shares (grey, dotted), by subgroup. In the bottom panel, the dashed diagonal marks perfect agreement between estimated probabilities and observed proportions. The area of each circle is proportional to the number of records in the bin.
Figure 6 : County-level subgroup-specific estimates of the share in the top quartile of the 1900 socioeconomic index. Panels give signed bias (left) and root-mean-square error (right), both in percentage points. Red bars are coarse prior ℓ BISG, blue bars oracle ℓ BISG, and grey bars recovered geography only.
Coverage
Precision
Sect
Registrants
Ground truth
Estimate
Ground truth
Estimate
Shia
1,096,570
0.72
0.72
0.48
0.48
Sunni
991,158
0.56
0.58
0.35
0.36
Maronite
724,231
0.75
0.77
0.37
0.38
Roman Orthodox
254,712
0.64
0.68
0.11
0.11
Druze
201,756
0.80
0.83
0.11
0.12
Table 3 : Quality of the seven LLM sect lists on the Lebanese voter registry. “Coverage” is the share of a sect’s registrants carrying a surname on that sect’s list, and “Precision” the share of the registrants carrying a list surname who belong to the sect. Lists contain up to 1,000 unique surnames.
Figure 7 : Precision and recall for ℓ BISG and geography alone for the seven principal Lebanese sects, by confession (rows) and sect (columns), under the locality-level prior ( n=1,550 ; ℓ BISG red solid, geography grey dotted) and the district-level prior ( n=29 ; ℓ BISG blue dashed, geography dark grey dot-dashed).
Figure 8 : Calibration of estimated individual sect probabilities based on ℓ BISG (LLM lists) for the seven principal Lebanese sects, by confession (rows) and sect (columns). The dashed diagonal marks perfect agreement between estimated probabilities and observed proportions. The area of each circle is proportional to the number of registrants in the bin.
Figure 9 : Share of each sect born in 1985 or later, estimated by district. Rows give signed bias and root-mean-square error against the enumerated sect, both in ppts; red bars are ℓ BISG and grey bars geography alone. Sects are ordered by their share of the roll.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A1 : Calibration of the raw LLM list score fr=P(Lr=1∣E) used directly as a race probability (blue, dashed) against ℓ BISG (red, solid), by race, on the Florida registrants. The dashed diagonal indicates perfect agreement between estimated probabilities and observed proportions. The area of each circle is proportional to the number of registrants in the bin.
Figure A2 : Surname embeddings (E5), Florida. A two-dimensional UMAP projection colored by dominant group (i.e., argmaxrfr ), with one exemplar label per group.
Figure A3 : The 428 Asian-dominant Florida surnames, colored by country-of-origin list. Grey points are on no origin list.
Figure A4 : The label-free criterion Q(K) for choosing the number of clusters on the one million record Florida and North Carolina samples. Panels facet state and name field; colors distinguish LLM and Census lists. Higher values are preferred. Open circles mark the selected value of K . The candidate grid ranges from 10 to 20,000 .
Figure A5 : The criterion Q(K) of equation ( 6 ) and the ground truth criterion for choosing the number of clusters K , by sample size, for the Florida surname model with the LLM lists. The ground truth is the PR-AUC (dashed) and the Brier score (dotted) against self-reported race, averaged over the four groups. The Brier score is reversed so that higher values are preferred for every curve. Each curve is rescaled to the unit interval within its panel, open circles denote the largest value of each curve, and the vertical line marks the selected K .
Figure A6 : Discrimination when the clusters are formed by K-means on the raw E5 surname embeddings (no lists) against K-means on the list scores. Florida sample with the block-level prior, by number of clusters K . Top row: recall at a precision of 0.5 . Bottom row: average precision.
Figure A7 : As Figure A6 , for the seven principal Lebanese sects, with the multilingual surname embeddings of Section 5.3 and the locality-level prior.
Group
Descriptions
Chinese
“Cantonese families from the Pearl River Delta settling in California and the American West”
“Taishanese and Sze Yup laborers migrating to the sugar plantations of the Hawaiian Islands”
“Fujianese and Hokkien merchants and sailors establishing communities in Atlantic port cities like New York”
Japanese
“Japanese immigrants from Hiroshima and Yamaguchi prefectures settling in Hawaii”
“Japanese immigrants from Kyushu and Okinawa settling in the United States and Hawaii”
“Japanese immigrants from Central and Northern Honshu settling in the Pacific Northwest and California”
Appendix
Table A1 : Model-written descriptions for the historical census lists.
Prompt stem
Completions
Lebanese Shia Muslim families from [x]
southern villages (Nabatieh, Tyre rural areas — ordinary families, NOT famous religious or political figures)
Baalbek-Hermel and the northern Beqaa (ordinary villagers)
the Beirut southern suburbs (Dahieh) — everyday families
Lebanese Shia Muslim families from [x] (ordinary families, not famous people)
the Nabatieh and Bint Jbeil districts
Beirut’s southern suburbs
the Baalbek-Hermel countryside
Appendix
Table A2 : Group descriptions used in the Lebanese sect prompts.
White
Black
Hispanic
Asian
τ
Shared
Cov.
Prec.
Cov.
Prec.
Cov.
Prec.
Cov.
Prec.
0.20
7.0%
0.28
0.79
0.51
0.37
0.50
0.82
0.42
0.60
0.30
0.04%
0.28
0.79
0.33
0.46
0.50
0.82
0.42
0.64
0.40
0.00%
0.28
0.79
0.17
0.62
0.50
0.82
0.41
0.65
0.50
0.00%
0.28
0.79
0.11
0.75
0.50
0.83
0.40
0.73
0.60
0.00%
0.28
0.79
0.08
0.85
0.49
0.83
0.39
0.74
Appendix
Table A3 : Sensitivity of the Census-based lists to the probability floor τ , where a surname enters the list for group r when Pr(R=r∣S=s)≥τ . Coverage and precision are defined as in Table 1 , computed on the Florida file. “Shared” is the share of registrants carrying a surname on more than one of the four lists. The floor used in the paper is 0.30 .
Figure A8 : Discrimination lift over geography alone, in ROC-AUC and PR-AUC, against the number of surnames per group. Calculated for the Florida (top row) and North Carolina (bottom row) voter files with published Census tract shares as the geographic baseline, for the Census and LLM lists of Table 1 . We use separate binary membership targets for each list. Standard BISG is included as a flat reference.
Figure A9 : List quality on the Florida voter file. Each cell is πr,r′ , the rate at which members of group r (rows) carry a name on list r′ (columns), for the Census lists (top) and the LLM lists (bottom). The left column computes Π from the self-reported race (i.e., ground-truth). The right column estimates Π from the tract-level list hit rates and the published tract race shares, with no individual race labels. The number above each panel is D(Π) .
Figure A10 : County-level race-specific estimates of the Democratic share of registrants with the LLM lists (red), the Census lists (green), and standard BISG/BIFSG (blue), otherwise as Figure 4 .