cs.CLOct 5, 2026

Probabilistic Race and Ethnicity Prediction Using Group-Specific Name Lists

Authors: Kyla Chasalow, Noah Dasanaike, Kosuke Imai

Organizations: PhD Candidate, Department of Statistics, Harvard University. · PhD Candidate, Department of Government, Harvard University. · Edith and Benjamin Geisinger Professor, Department of Government and Department of Statistics, Harvard University. 1737 Cambridge Street, Institute for Quantitative Social Science, Cambridge MA 02138.

Abstract

Statistically valid estimation of racial and ethnic disparities often requires inferring the probability that an individual belongs to a particular racial or ethnic group given only their name and geographic location. The standard approach, Bayesian Improved Surname Geocoding (BISG), relies on group population frequencies for each name. Although the U.S. Census Bureau provides such information for common names and a limited set of racial categories, comparable data do not exist for many racial and ethnic groups and are rarely available outside the U.S. We propose the list-powered BISG (ℓ\ellBISG) method, which can be used to derive calibrated group probabilities from group-specific name lists. These lists may be compiled based on expert knowledge or generated synthetically using large language models (LLMs), and thus may be subject to unknown biases. Representing names as embeddings, we treat list membership as a proxy prediction task and apply a correction based on proximal inference to recover the target group probabilities. We validate the method on U.S. voter files with self-reported race, on the full-count 1900 U.S. Census, and on the Lebanese voter registry. We find that LLM-generated name lists yield accurate and well-calibrated probabilities as well as precise disparity estimates comparable to those obtained using methods that require name-race data. Thus, ℓ\ellBISG substantially broadens the applicability of probabilistic race and ethnicity prediction to settings where name-race data are unavailable.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Using Embedding Models to Improve Probabilistic Race Prediction

    Apr 24, 2026Noah Dasanaike, Kosuke ImaiLineages

  2. Productionized Fairness Measurement Under Privacy Constraints

    Jun 25, 2026Osonde A. Osoba, Yuzi He, Saikrishna Badrinarayanan +3Algorithmic FairnessMeasures

  3. Tracing the Latent Threads: A Mechanistic Study of How LLMs Represent and Operationalize Race and Ethnicity Cues

    Jan 19, 2026Shiyue Hu, Ruizhe Li, Yanjun GaoLarge Language Model BiasDisparities