How Does "English (US)" Become the Default? Triangulating Structural Bias Towards American English Across the LLM Pipeline
Organizations: Department of Computing Science University of Alberta
Abstract
Large language models (LLMs) are increasingly embedded in educational, professional, and public infrastructure, yet widely used platforms expose "English (US)" as a primary English setting despite the global diversity of English. We ask: How does "English (US)" become the default? We study this question as structural bias, examining how geopolitical histories of data curation, digital dominance, and linguistic standardization intersect with the LLM development pipeline. Using British English as a controlled reference, we construct a curated resource of 1,813 matched American English (AmE)--British English (BrE) variants and introduce DiAlign, a dynamic, training-free method for estimating regional alignment from distributional evidence. We triangulate the AmE preference across data exposure --> representation --> generation, jointly examining pretraining and post-training data, tokenizer behavior and provenance, model prediction cost, and generated language across developer countries, prompt conditions, domains and sources, linguistic categories, and registers. AmE is consistently favored across all six audited pretraining corpora and 21 post-training datasets, is generally represented more compactly by tokenizers, and receives lower prediction cost. It also remains the dominant generation default under neutral English prompting; British-English prompting shifts this preference toward BrE but does not consistently eliminate the AmE default. To our knowledge, this is the first rigorous pipeline-wide study of structural bias across major phases of LLM development. Our findings show that contemporary LLMs privilege AmE as the de facto norm, raising concerns about linguistic homogenization, epistemic injustice, and inequity in global AI deployment, while providing a rigorous basis for targeted component-level intervention.
Figures & tables
| Lexical Variants | DiAlign (broader variants) | |||||||||
| Data Source | Document Type | Documents | Tokens | Orthographic | Vocabulary | % of | ||||
| (millions) | (billions) | AmE ( ) | ( ) | AmE ( ) | ( ) | AmE ( ) | ( ) | |||
| Book Corpus ( 2015 ) | books | 74 | 1.28 | 86.81 | 13.19 | 75.00 | 25.00 | 79.90 [0.79] | 20.10 [0.76] | |
| Wikipedia ( 2024 ) | encyclopedic | 6.4 | 4.3 | 72.94 | 27.06 | 61.43 | 38.57 | 72.18 [0.80] | 27.82 [0.78] | |
| Common Crawl (C4) ( 2020 ) | web pages | 365 | 156 | 75.12 | 24.88 | 67.00 | 33.00 | 73.67 [0.81] | 26.33 [0.81] | |
| Falcon RefinedWeb ( 2023 ) | web pages | 968 | 600 | 77.34 | 22.66 | 68.35 | 31.65 | 74.04 [0.82] | 25.96 [0.75] | |
| Post-training stage | Hits | AmE (%) |
| Instruction tuning / SFT | 16,348,769 | 75.95 |
| Instruction tuning / chat | 1,204,637 | 84.35 |
| Instruction tuning / hum feedback | 61,030 | 83.08 |
| Instruction tuning / task mixture | 6,578,816 | 71.77 |
| RLHF / preference data | 4,168,689 | 80.47 |
| RLHF / reward-model feedback | 77,332 | 84.72 |
| Orthographic | Vocabulary | ||||||||||
| Tokenizers | Developer Country | Model Access | Vocab Size | AmE ( ) | ( ) | AmE ( ) | ( ) | ||||
| GPT-4 | USA ( ) | 100K | 2.73 | 2.86 | 4.76 % | 2.27 | 2.64 | 16.30 % | |||
| GPT-4o | USA ( ) | 200K | 2.65 | 2.77 | 4.53 % | 2.21 | 2.57 | 16.29 % | |||
| Llama-3.3-70B | USA ( ) | 128K | 2.72 | 2.85 | 4.78 % | 2.27 | 2.63 | 15.86 % | |||
| Gemma-3-27B | USA ( ) | 262K | 2.40 | 2.53 | 5.42 % | 2.02 | 2.35 | 16.34 % | |||
| DeepSeek-V3 | China ( ) | 128K | 2.71 | 2.80 | 3.32 % | 2.37 | 2.67 | 12.66 % | |||
| Natural Questions [ formal ] | ELI5 [ informal ] | ||||||
| LLMs | Developer Country | Default English ( ) | British English ( ) | Default English ( ) | British English ( ) | ||
| GPT-4o | USA ( ) | 79.00% [0.81] | 45.33% [0.77] | 77.00% [0.82] | 3 4.67% [0.78] | ||
| Gemini-2.0-flash | USA ( ) | 76.00% [0.83] | 51.00% [0.79] | 75.33% [0.82] | 42.33% [0.80] | ||
| Claude-3.7-sonnet | USA ( ) | 75.67% [0.85] | 4 2.33% [0.80] | 73.33% [0.86] | 37.67% [0.78] | ||
| Llama-3.3-70B | USA ( ) | 74.67% [0.82] | 47.33% [0.78] | 69.00% [0.79] | 30.00% [0.76] | ||
| Gemma-3-27B | USA ( ) | 69.33% [0.81] | 45.67% [0.78] | 6 8.33% [0.83] | 38.00% [0.78] | ||
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
| Category | Difference Type | % of Pairs | Examples | ||
| ends in “-or” (AmE) vs. “-our” (BrE) | 2.26% | color ( ) vs. colour ( ) | labor ( ) vs. labour ( ) | ||
| ends in “-ize” (AmE) vs. “-ise” (BrE) | 11.58% | organize ( ) vs. organise ( ) | realize ( ) vs. realise ( ) | ||
| ends in “-er” (AmE) vs. “-re” (BrE) | 1.65% | center ( ) vs. centre ( ) | liter ( ) vs. litre ( ) | ||
| ends in “-og” (AmE) vs. “-ogue” (BrE) | 0.55% | dialog ( ) vs. dialogue ( ) | catalog ( ) vs. catalogue ( ) | ||
| ends in “-ense” (AmE) vs. “-ence” (BrE) | 0.22% | defense ( ) vs. defence ( ) | pretense ( ) vs. pretence ( ) | ||
| “e” (AmE) vs. “ae” (BrE) | 4.03% | esthetic ( ) vs. aesthetic ( ) | pediatric ( ) vs. paediatric ( ) | ||
| Level | American English (AmE) ( ) | Reference: British English (BrE) ( ) |
| Orthography | color, center, organize, traveling, favorite, check, jewelry, program, catalog, defense | colour, centre, organise, travelling, favourite, cheque, jewellery, programme, catalogue, defence |
| Vocabulary | cell phone, crosswalk, downtown, train station, parking lot, ZIP code, vacation, apartment, elevator, line, sidewalk, flashlight, soda, popsicle | mobile phone, zebra crossing, city centre, railway station, car park, postcode, holiday, flat, lift, queue, pavement, torch, fizzy drink, ice lolly |
| Grammar/usage | I just ate ; on the weekend ; Do you have a pen? ; The team is winning ; different than | I’ve just eaten ; at the weekend ; Have you got a pen? ; The team are winning ; different from |
| Conventions | double quotes; punctuation inside quotes; December 31, 2024 ; first floor ; 11:15 PM ; F | single quotes; punctuation outside quotes; 31 December 2024 ; ground floor ; 11.15 pm ; C |
| Ablations | Accuracy | Precision | Recall | F1 Score | Avg. Conf. (AmE) | Avg. Conf. (BrE) |
| DiAlign (final) | 93.18 | 90.67 | 96.25 | 93.38 | 0.84 | 0.91 |
| – w/o Divergence Weight (DW) | 92.25 | 89.29 | 95.98 | 92.52 | 0.77 | 0.85 |
| – w/o Boosting Factor (BF) | 92.25 | 89.29 | 95.98 | 92.52 | 0.82 | 0.89 |
| – w/o Both (DW + BF) | 91.31 | 88.13 | 95.45 | 91.65 | 0.75 | 0.83 |
| Reference period | Accuracy | Precision | Recall | F1 |
| 1950–2022 | 93.18 | 90.67 | 96.25 | 93.38 |
| 2000–2022 | 92.91 | 90.42 | 95.98 | 93.12 |
| Dataset | Stage | Rows | Hits | Weighted AmE (%) | Orth. AmE (%) | Vocab. AmE (%) | HF |
| Alpaca Cleaned ( Taori et al., 2023 ) | SFT | 51,760 | 74,991 | 85.35 | 94.84 | 80.05 | HF |
| Databricks Dolly 15k ( Conover et al., 2023 ) | SFT | 15,011 | 19,997 | 65.29 | 78.09 | 67.65 | HF |
| OpenAssistant Conversations v2 ( Köpf et al., 2023 ) | SFT | 61,278 | 61,030 | 83.08 | 92.56 | 76.85 | HF |
| UltraChat 200k ( Ding et al., 2023 ) | SFT | 207,865 | 2,442,647 | 84.23 | 90.32 | 75.68 | HF |
| OpenHermes 2.5 ( Teknium, 2023 ) | SFT | 1,001,551 | 2,517,105 | 81.81 | 85.78 | 74.01 | HF |
| WizardLM Evol-Instruct v2 ( Xu et al., 2024 ) | SFT | 143,000 | 634,919 | 89.72 | 95.20 | 80.33 | HF |
| Tokenizer | Developer Country | Tokenizer lineage | Independent | Interpretation |
| GPT-4 | USA ( ) | OpenAI cl100k_base | Yes | Independent US baseline |
| GPT-4o | USA ( ) | OpenAI o200k_base | Yes | Independent US baseline |
| Llama-3.3-70B | USA ( ) | Extended from US-origin tiktoken-style vocabulary | No | Derived; excluded from independent country-level evidence |
| Gemma-3-27B | USA ( ) | Gemma/Gemini-family vocabulary | Yes | Independent US tokenizer |
| DeepSeek-V3 | China ( ) | DeepSeek BPE | Yes | Independent non-US tokenizer |
| Mistral-Small-24B | France ( ) | Mistral Tekken family | Yes | Independent non-US tokenizer |
| Tokenizer compared with GPT-4 | AmE | BrE | Validation | Generated | All |
| StableLM-2-1.6B | 100.00 | 100.00 | 80.80 | 81.17 | 91.87 |
| Llama-3.3-70B | 99.17 | 99.06 | 99.87 | 99.83 | 99.43 |
| GPT-4o | 74.85 | 77.05 | 40.27 | 29.00 | 58.58 |
| DeepSeek-V3 | 67.02 | 70.22 | 47.80 | 25.08 | 55.42 |
| Mistral-Small-24B | 65.47 | 66.91 | 26.47 | 10.00 | 46.11 |
| Falcon3-7B | 58.58 | 58.91 | 11.93 | 5.17 | 37.48 |
| Family | Stage | Rows | Gap / token | 95% CI | higher |
| Llama | Base | 1,008 | 0.189 | [0.185, 0.193] | 100.0% |
| Llama | Post | 1,008 | 0.264 | [0.262, 0.267] | 100.0% |
| StableLM | Base | 2,350 | 0.210 | [0.204, 0.214] | 100.0% |
| StableLM | Post | 2,350 | 0.400 | [0.389, 0.411] | 100.0% |
| Ministral | Base | 1,068 | 0.213 | [0.208, 0.218] | 100.0% |
| Ministral | Post | 1,068 | 0.234 | [0.229, 0.239] | 100.0% |
| Family | Stage | Hugging Face checkpoint |
| Gemma | Base | google/gemma-3-1b-pt |
| Gemma | Post-trained | google/gemma-3-1b-it |
| Ministral | Base | mistralai/Ministral-3-3B-Base-2512 |
| Ministral | Post-trained | mistralai/Ministral-3-3B-Instruct-2512 |
| Llama | Base | meta-llama/Llama-3.1-8B |
| Llama | Post-trained | meta-llama/Llama-3.1-8B-Instruct |
| Layer | Comparison | Models | Mean cosine |
| Embedding | – counterfactual | 10 | 0.902 |
| Embedding | Character perturbation | 10 | 0.824 |
| Embedding | Unrelated control | 10 | 0.504 |
| Middle | – counterfactual | 10 | 1.000 |
| Middle | Character perturbation | 10 | 0.999 |
| Middle | Unrelated control | 10 | 0.997 |
| Natural Questions (NQ) [formal] | ELI5 [informal] | ||||||||||||
| Default English ( ) | British English ( ) | Default English ( ) | British English ( ) | ||||||||||
| LLMs | Model Access | #Words Avg. [SD] | Range [min-max] | #Words Avg. [SD] | Range [min-max] | #Words Avg. [SD] | Range [min-max] | #Words Avg. [SD] | Range [min-max] | ||||
| GPT-4o | 50.19 [1.75] | [46–57] | 50.32 [1.61] | [45–56] | 50.41 [1.60] | [47–55] | 50.35 [1.53] | [46–55] | |||||
| Gemini-2.0-flash | 52.74 [2.71] | [46–60] | 51.92 [2.89] | [45–61] | 53.10 [2.62] | [47–61] | 52.82 [2.95] | [44–60] | |||||
| Claude-3.7-sonnet | 45.25 [2.29] | [39–52] | 45.49 [2.35] | [39–57] | 47.70 [2.45] | [42–56] | 47.38 [2.32] | [41–54] | |||||
| Llama-3.3-70B | 41.68 [7.12] | [16–50] | 41.47 [7.44] | [18–50] | 41.33 [7.83] | [21–50] | 40.10 [8.48] | [18–51] | |||||
| Condition | Length bin | Avg. words | DiAlign AmE (%) |
| Default | 1–50 | 44.08 | 67.64 |
| Default | 51–75 | 55.82 | 67.33 |
| Default | 76–100 | 86.70 | 64.31 |
| Default | 101–150 | 116.03 | 62.54 |
| Default | 151+ | 225.60 | 61.01 |
| British-English steer | 1–50 | 43.62 | 47.12 |
| Source | Title ( linked ) | Description |
| Wikipedia | American and British English spelling differences | A widely cited reference outlining systematic orthographic differences between American and British English. The page provides examples of variant spellings (e.g., color vs. colour ), historical background, and explanations of regional conventions. It served as one of the authentic linguistic resources for curating consistent one-to-one variant pairs in our lexicon. |
| ThoughtCo. | American English to British English Vocabulary | A curated reference list of American and British English vocabulary differences, created by experienced educators and subject experts. Provides reliable lexical contrasts in an accessible format, supporting the construction of our AmE–BrE lexicon. |
| Research Article | Mapping the Americanization of English in Space and Time | An empirical study tracing how American English variants spread globally across regions and over time. Offers quantitative evidence of AmE–BrE lexical contrasts, providing authoritative grounding for the curated variant pairs in our unified lexicon. |
| IELTS | British vs. American English in the IELTS Test: Key Differences | An official IELTS guide highlighting key vocabulary, spelling, and grammar differences between AmE and BrE. The resource systematically documents contrasts across domains such as food, school, homes, and grammar, making it a practical reference for understanding standardized English variations. |
| Grammarly | How to Select Your English Dialect | A practical guide from Grammarly explaining how to switch between English dialects in writing tools, highlighting spelling, vocabulary, and usage variations (AmE vs BrE). Because it enumerates common dialectal choices in real writing, it serves as a useful supplementary resource for identifying variant pairs. |
| SpellZone | Sixty American English Words and their British English Counterparts | SpellZone provides a practical reference list of 60 common AmE–BrE word pairs, illustrating clear lexical contrasts in spelling and vocabulary. The resource highlights straightforward one-to-one mappings useful for systematic dialectal analysis. |
| Category | British English (BrE) | American English (AmE) |
| o vs. ou | colour, honour, behaviour | color, honor, behavior |
| -re vs. -er endings | centre, fibre, theatre | center, fiber, theater |
| -ise vs. -ize endings | recognise, authorise | recognize, authorize |
| -yse vs. -yze endings | analyse, paralyse, catalyse | analyze, paralyze, catalyze |
| Single vs. double l (inflection) | travelled, counselled | traveled, counseled |
| -ll + -ly suffix | skilfully, wilfully | skillfully, willfully |
| Category | British English (BrE) | American English (AmE) |
| Preposition before days | She resigned on Thursday. | She resigned Thursday. |
| Street naming | in the High Street | on Main Street |
| Transitivity (protest) | protest against discrimination | protest discrimination |
| Ditransitives (write) | write to me | write me |
| Meeting collocation | meet the team | meet with the team |
| Transport & wayfinding |