The "10th Juror": Open-Set Standpoint Screening for Bureaucratic Bias Detection
Organizations: Sydney Smart Technology College, Northeastern University, China · Taiyuan University of Technology · School of Resources, Environment and Materials, Guangxi University · School of Computer and Communication Engineering, Northeastern University at Qinhuangdao, China
Abstract
Presupposing the boundaries of bias is itself a form of bias. We study closed-loop bias governance for Dutch government documents, where a system must detect biased language, ground decisions in legal and contextual evidence, rewrite problematic sentences when intervention is warranted, and verify that the rewrite mitigates harm without distorting meaning. Existing methods face three challenges: (i) discriminative classifiers capture surface regularities but lack normative grounding; (ii) zero-shot LLMs often adopt generic viewpoints and over-flag ambiguous administrative language; and (iii) fixed taxonomies inherit the Closed-World Assumption, missing emerging local targets. We propose MARS-Gov, a standpoint-aware multi-agent framework that combines legal retrieval, open-set target screening, specialized jurors, conservative routing, and rewrite verification. When screening finds an uncovered group, MARS-Gov instantiates a dynamic "10th juror" to deliberate outside the fixed panel. On DGDB, MARS-Gov sets a new SOTA with 0.880 F1, outperforming the strongest zero-shot LLM detector by 20.2 points (29.8% relative) and the best supervised Dutch encoder by 6.8 points, while reducing unnecessary interventions to 2.5%. Leave-One-Category-Out (LOCO) evaluation recovers held-out categories with 85.1% Correct@1 and 93.8% Correct@3.
Figures & tables
| Prec. | Rec. | F1 | F1 | Drop | OGR | BMR | SF | BGS | |
| Fine-tuned | |||||||||
| BERTje de Vries et al. (2019) | 0.842 .017 | 0.784 .021 | 0.812 .014 | 0.554 .032 | -31.8% 1.62 | 3.1% 0.51 | - | - | - |
| RobBERT Delobelle et al. (2020) | 0.839 .019 | 0.784 .018 | 0.811 .013 | 0.499 .034 | -38.5% 1.94 | 3.6% 0.62 | - | - | - |
| I. Prompted LLMs (Zero-/Few-shot) | |||||||||
| GPT-4o | 0.514 .025 | 0.882 .021 | 0.649 .017 | 0.621 .024 | -4.3% 1.07 | 16.4% 0.83 | 0.872 .016 | 0.821 .021 | 0.631 .021 |
| Grok-3 | 0.484 .031 | 0.920 .017 | 0.634 .019 | 0.598 .029 | -5.7% 1.18 | 19.1% 0.94 | 0.864 .019 | 0.808 .024 | 0.642 .022 |
| Config. | Metric | B1 | B2 | B3 | B4 | B5 | B6 | B7 | B8 | B9 |
|---|---|---|---|---|---|---|---|---|---|---|
| (Disability) | (Migration) | (Colonialism) | (Religion) | (Gender) | (LGBTQ+) | (Culture) | (Education/Class) | (Welfare) | ||
| w/o Scout | Rec. | 0.794 .027 | 0.814 .024 | 0.778 .031 | 0.776 .028 | 0.821 .023 | 0.788 .026 | 0.807 .024 | 0.764 .029 | 0.812 .022 |
| F1 | 0.804 .021 | 0.826 .018 | 0.794 .025 | 0.789 .022 | 0.831 .017 | 0.806 .019 | 0.813 .018 | 0.782 .023 | 0.819 .017 | |
| w/ Scout (Ours) | Rec. | 0.882 .018 | 0.901 .015 | 0.874 .021 | 0.871 .019 | 0.894 .014 | 0.881 .018 | 0.889 .016 | 0.866 .021 | 0.896 .015 |
| F1 | 0.873 .015 | 0.891 .012 | 0.867 .017 | 0.864 .015 | 0.886 .011 | 0.872 .014 | 0.881 .013 | 0.854 .017 | 0.884 .012 | |
| Correct@1 (%) | 83.4 2.4 | 87.6 1.9 | 82.7 2.7 | 86.5 2.1 | 88.9 1.7 | 84.3 2.2 | 81.6 2.6 | 83.2 2.5 | 87.4 1.8 |
| Configuration | Detection Quality | Governance | Efficiency | ||
|---|---|---|---|---|---|
| Prec. | Rec. | F1 | BGS | Tokens/Doc | |
| Full MARS-Gov | 0.865 .015 | 0.895 .013 | 0.880 .012 | 0.808 .013 | 2,450 58 |
| w/o Verify Loop | 0.865 .015 | 0.895 .013 | 0.880 .012 | 0.646 .019 ( 20.0%) ∗∗ | 2,100 49 |
| w/o Reasoning | 0.842 .019 | 0.870 .017 | 0.856 .016 | 0.765 .021 ( 5.3%) ∗ | 2,380 62 |
| w/o NEC | 0.802 .021 | 0.847 .019 | 0.824 .018 | 0.710 .022 ( 12.1%) ∗∗ | 2,250 54 |
| w/o Scout (10th) | 0.851 .017 | 0.840 .019 | 0.845 .015 | 0.760 .018 ( 5.9%) ∗ | 2,150 44 |
| Configuration | Calls | Lat. (s) | Clear% | $/1k |
|---|---|---|---|---|
| Baselines (Gemini 2.5 Pro) | ||||
| Zero-shot LLM | 1.0 | 1.8 (0.3) | — | 0.85 |
| Standard Debate | 4.0 | 7.2 (0.9) | — | 4.62 |
| SA-RAG | 1.0 | 3.4 (0.5) | — | 2.46 |
| MARS-Gov (with ARG) | ||||
| GPT-4o | 12.6 (0.6) | 20.5 (2.3) | 64.2 (1.7) | 10.04 |
| Configuration | DALC (Dutch, Social Media) | SBIC (English, Social Media) | KoBBQ (Korean, General) | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Prec. | Rec. | F1 | Prec. | Rec. | F1 | Prec. | Rec. | F1 | |
| Zero-shot LLMs | 0.513 .028 | 0.873 .021 | 0.646 .018 | 0.486 .033 | 0.844 .023 | 0.617 .019 | 0.451 .036 | 0.803 .027 | 0.578 .024 |
| Standard Debate | 0.607 .025 | 0.851 .020 | 0.709 .019 | 0.583 .027 | 0.819 .021 | 0.681 .021 | 0.553 .029 | 0.811 .024 | 0.658 .023 |
| SA-RAG | 0.653 .021 | 0.862 .018 | 0.743 .016 | 0.626 .023 | 0.836 .019 | 0.716 .018 | 0.593 .027 | 0.821 .021 | 0.689 .019 |
| Ours: MARS-Gov | 0.823 .015 | 0.886 .013 | 0.853 .012 | 0.803 .018 | 0.847 .016 | 0.824 .015 | 0.776 .021 | 0.829 .018 | 0.802 .017 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Statistics | DGDB | DGDB-Unseen |
|---|---|---|
| Language | Dutch | Dutch |
| Domain | Gov. docs | Gov. docs |
| Unit | Sentence | Sentence |
| Corpus size | 3,747 | DGDB-derived |
| Evaluation size | Test split | Held-out subset |
| Construction | Original split | Bias-term substitution |
| Statistics | DALC | SBIC | KoBBQ |
|---|---|---|---|
| Language | Dutch | English | Korean |
| Domain | Social | QA | |
| Unit | Tweet | Post / tuple | MCQ |
| Corpus size | 8,156 | 44,671 / 147,139 | 76,048 |
| Evaluation size | 1,000 | 1,000 | 1,000 |
| Evaluation balance | 500 / 500 | 500 / 500 | 500 / 500 |
| Model Family | Model Name | Standard Zero-shot | Ours: MARS-Gov | Governance | ||||
| Prec | Rec | F1 | Prec | Rec | F1 | BGS | ||
| Frontier SOTA | Gemini 2.5 Pro | 0.524 .024 | 0.907 .018 | 0.664 .016 | 0.865 .015 | 0.895 .013 | 0.880 .012 | 0.808 .013 |
| DeepSeek-R1 | 0.549 .023 | 0.887 .019 | 0.678 .015 | 0.872 .016 | 0.884 .014 | 0.878 .013 | 0.806 .015 | |
| Claude 3.5 Sonnet | 0.562 .022 | 0.872 .020 | 0.683 .016 | 0.860 .017 | 0.882 .015 | 0.871 .013 | 0.787 .016 | |
| Grok-3 | 0.484 .031 | 0.920 .017 | 0.634 .019 | 0.836 .019 | 0.910 .013 | 0.871 .014 | 0.779 .018 | |
| GPT-4o | 0.514 .025 | 0.882 .021 | 0.649 .017 | 0.844 .017 | 0.863 .015 | 0.853 .013 | 0.757 .014 | |
| Dimension | DGDB NEC retrieval |
|---|---|
| Primary Objective | Detecting implicit bias in official policy texts. |
| Core Domain | Legal and political administrative text ( high-stakes ). |
| Inventory | Dutch constitutional and equal-treatment law; antidiscrimination policy; accessibility and election communication guidance; online-discrimination, institutional-racism, emancipation, and inclusive-language resources; DGDB reasoning entries. |
| Selection Rationale | Grounds the review in legal norms while covering administrative, accessibility, and inclusive-language contexts that commonly shape implicit bias. |