Organizations: Sydney Smart Technology College, Northeastern University, China · Taiyuan University of Technology · School of Resources, Environment and Materials, Guangxi University · School of Computer and Communication Engineering, Northeastern University at Qinhuangdao, China
Presupposing the boundaries of bias is itself a form of bias. We study closed-loop bias governance for Dutch government documents, where a system must detect biased language, ground decisions in legal and contextual evidence, rewrite problematic sentences when intervention is warranted, and verify that the rewrite mitigates harm without distorting meaning. Existing methods face three challenges: (i) discriminative classifiers capture surface regularities but lack normative grounding; (ii) zero-shot LLMs often adopt generic viewpoints and over-flag ambiguous administrative language; and (iii) fixed taxonomies inherit the Closed-World Assumption, missing emerging local targets. We propose MARS-Gov, a standpoint-aware multi-agent framework that combines legal retrieval, open-set target screening, specialized jurors, conservative routing, and rewrite verification. When screening finds an uncovered group, MARS-Gov instantiates a dynamic "10th juror" to deliberate outside the fixed panel. On DGDB, MARS-Gov sets a new SOTA with 0.880 F1, outperforming the strongest zero-shot LLM detector by 20.2 points (29.8% relative) and the best supervised Dutch encoder by 6.8 points, while reducing unnecessary interventions to 2.5%. Leave-One-Category-Out (LOCO) evaluation recovers held-out categories with 85.1% Correct@1 and 93.8% Correct@3.
Figures & tables
Figure 1: Closed-set detection versus MARS-Gov open-set screening. Fixed taxonomies may miss local targets or force them into a known class. MARS-Gov retrieves legal evidence, detects standpoint gaps, adds a dynamic 10th juror when needed, and sends grounded reports to the Judge for an auditable verdict and minimal rewrite.
Figure 2: Overview of MARS-Gov. The framework retrieves normative evidence, gathers juror risk reports, routes only risky cases to a Judge, and verifies any minimal rewrite. Separating evidence, routing, adjudication, and restoration keeps the decision trace inspectable.
Figure 3: Three-layer retrieval in NEC.
Prec.
Rec.
F1
F1 OOD
Δ Drop
OGR ↓
BMR
SF
BGS
Fine-tuned
BERTje de Vries et al. (2019)
0.842 ± .017
0.784 ± .021
0.812 ± .014
0.554 ± .032
-31.8% ± 1.62
3.1% ± 0.51
-
-
-
RobBERT Delobelle et al. (2020)
0.839 ± .019
0.784 ± .018
0.811 ± .013
0.499 ± .034
-38.5% ± 1.94
3.6% ± 0.62
-
-
-
I. Prompted LLMs (Zero-/Few-shot)
GPT-4o
0.514 ± .025
0.882 ± .021
0.649 ± .017
0.621 ± .024
-4.3% ± 1.07
16.4% ± 0.83
0.872 ± .016
0.821 ± .021
0.631 ± .021
Grok-3
0.484 ± .031
0.920 ± .017
0.634 ± .019
0.598 ± .029
-5.7% ± 1.18
19.1% ± 0.94
0.864 ± .019
0.808 ± .024
0.642 ± .022
Table 1: Main DGDB results. Darker green marks stronger performance; darker red marks larger Δ Drop or OGR. F1 OOD is DGDB-Unseen F1; Δ Drop is the relative drop from DGDB F1. Single-Call was evaluated in-domain only. MARS-Gov gives the best mean F1/BGS with low governance intrusion.
Config.
Metric
B1
B2
B3
B4
B5
B6
B7
B8
B9
(Disability)
(Migration)
(Colonialism)
(Religion)
(Gender)
(LGBTQ+)
(Culture)
(Education/Class)
(Welfare)
w/o Scout
Rec.
0.794 ± .027
0.814 ± .024
0.778 ± .031
0.776 ± .028
0.821 ± .023
0.788 ± .026
0.807 ± .024
0.764 ± .029
0.812 ± .022
F1
0.804 ± .021
0.826 ± .018
0.794 ± .025
0.789 ± .022
0.831 ± .017
0.806 ± .019
0.813 ± .018
0.782 ± .023
0.819 ± .017
w/ Scout (Ours)
Rec.
0.882 ± .018
0.901 ± .015
0.874 ± .021
0.871 ± .019
0.894 ± .014
0.881 ± .018
0.889 ± .016
0.866 ± .021
0.896 ± .015
F1
0.873 ± .015
0.891 ± .012
0.867 ± .017
0.864 ± .015
0.886 ± .011
0.872 ± .014
0.881 ± .013
0.854 ± .017
0.884 ± .012
Correct@1 (%)
83.4 ± 2.4
87.6 ± 1.9
82.7 ± 2.7
86.5 ± 2.1
88.9 ± 1.7
84.3 ± 2.2
81.6 ± 2.6
83.2 ± 2.5
87.4 ± 1.8
Table 2: Leave-One-Category-Out (LOCO) evaluation on Gemini 2.5 Pro. Best F1 scores are bold. Significance: all w/ Scout gains over w/o Scout remain significant after Bonferroni correction over 9 categories (two-sided paired t-test over 3 runs; max corrected p=2.4×10−2 ).
Configuration
Detection Quality
Governance
Efficiency
Prec.
Rec.
F1
BGS
Tokens/Doc
Full MARS-Gov
0.865 ± .015
0.895 ± .013
0.880 ± .012
0.808 ± .013
2,450 ± 58
w/o Verify Loop
0.865 ± .015
0.895 ± .013
0.880 ± .012
0.646 ± .019 ( ↓ 20.0%) ∗∗
2,100 ± 49
w/o Reasoning
0.842 ± .019
0.870 ± .017
0.856 ± .016
0.765 ± .021 ( ↓ 5.3%) ∗
2,380 ± 62
w/o NEC
0.802 ± .021
0.847 ± .019
0.824 ± .018
0.710 ± .022 ( ↓ 12.1%) ∗∗
2,250 ± 54
w/o Scout (10th)
0.851 ± .017
0.840 ± .019
0.845 ± .015
0.760 ± .018 ( ↓ 5.9%) ∗
2,150 ± 44
Table 3: Ablation on DGDB with Gemini 2.5 Pro, averaged over 3 runs. Significance: marked BGS drops vs. Full MARS-Gov remain significant after Bonferroni correction (two-sided paired t-test; ∗∗p<0.01 , ∗p<0.05 ; max corrected p=2.7×10−2 ). w/o ARG is cost-only.
Figure 4: Recall vs. Over-Governance Rate (OGR) trade-off as threshold τ increases.
Configuration
Calls
Lat. (s)
Clear%
$/1k
Baselines (Gemini 2.5 Pro)
Zero-shot LLM
1.0
1.8 (0.3)
—
0.85
Standard Debate
4.0
7.2 (0.9)
—
4.62
SA-RAG
1.0
3.4 (0.5)
—
2.46
MARS-Gov (with ARG)
GPT-4o
12.6 (0.6)
20.5 (2.3)
64.2 (1.7)
10.04
Table 4: Per-sentence inference cost on DGDB (1,000 sentences, 3 runs; mean with σ ). Calls counts all backbone calls; Clear% is the early-stop share.
Configuration
DALC (Dutch, Social Media)
SBIC (English, Social Media)
KoBBQ (Korean, General)
Prec.
Rec.
F1
Prec.
Rec.
F1
Prec.
Rec.
F1
Zero-shot LLMs
0.513 ± .028
0.873 ± .021
0.646 ± .018
0.486 ± .033
0.844 ± .023
0.617 ± .019
0.451 ± .036
0.803 ± .027
0.578 ± .024
Standard Debate
0.607 ± .025
0.851 ± .020
0.709 ± .019
0.583 ± .027
0.819 ± .021
0.681 ± .021
0.553 ± .029
0.811 ± .024
0.658 ± .023
SA-RAG
0.653 ± .021
0.862 ± .018
0.743 ± .016
0.626 ± .023
0.836 ± .019
0.716 ± .018
0.593 ± .027
0.821 ± .021
0.689 ± .019
Ours: MARS-Gov
0.823 ± .015
0.886 ± .013
0.853 ± .012
0.803 ± .018
0.847 ± .016
0.824 ± .015
0.776 ± .021
0.829 ± .018
0.802 ± .017
Table 5: Generalization on DALC, SBIC, and KoBBQ using Gemini 2.5 Pro. Results average 3 runs; best F1 scores are bold. Significance: MARS-Gov beats SA-RAG after Bonferroni correction over 3 datasets (two-sided paired t-test; max corrected p=1.1×10−2 ).
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Statistics
DGDB
DGDB-Unseen
Language
Dutch
Dutch
Domain
Gov. docs
Gov. docs
Unit
Sentence
Sentence
Corpus size
3,747
DGDB-derived
Evaluation size
Test split
Held-out subset
Construction
Original split
Bias-term substitution
Appendix
Table 6: Statistical overview of the DGDB-derived datasets.
Statistics
DALC
SBIC
KoBBQ
Language
Dutch
English
Korean
Domain
Twitter
Social
QA
Unit
Tweet
Post / tuple
MCQ
Corpus size
8,156
44,671 / 147,139
76,048
Evaluation size
1,000
1,000
1,000
Evaluation balance
500 / 500
500 / 500
500 / 500
Appendix
Table 7: Statistical overview of the external evaluation datasets.
Model Family
Model Name
Standard Zero-shot
Ours: MARS-Gov
Governance
Prec
Rec
F1
Prec
Rec
F1
BGS
Frontier SOTA
Gemini 2.5 Pro
0.524 ± .024
0.907 ± .018
0.664 ± .016
0.865 ± .015
0.895 ± .013
0.880 ± .012
0.808 ± .013
DeepSeek-R1
0.549 ± .023
0.887 ± .019
0.678 ± .015
0.872 ± .016
0.884 ± .014
0.878 ± .013
0.806 ± .015
Claude 3.5 Sonnet
0.562 ± .022
0.872 ± .020
0.683 ± .016
0.860 ± .017
0.882 ± .015
0.871 ± .013
0.787 ± .016
Grok-3
0.484 ± .031
0.920 ± .017
0.634 ± .019
0.836 ± .019
0.910 ± .013
0.871 ± .014
0.779 ± .018
GPT-4o
0.514 ± .025
0.882 ± .021
0.649 ± .017
0.844 ± .017
0.863 ± .015
0.853 ± .013
0.757 ± .014
Appendix
Table 8: Extended Leaderboard on DGDB (In-Domain). Models are ranked by their MARS-Gov F1 score. The red dashed line marks the performance threshold of the fine-tuned BERTje baseline ( F1=0.812 ). Observation: While MARS-Gov improves performance across model families (avg. Δ F1 +0.20 , larger for frontier backbones than for smaller ones), a minimum reasoning capability threshold (approx. GPT-4 Turbo level) is required to surpass supervised baselines.
Figure 5: Comparison of F1 scores between Native Dutch and Translated English prompts. Native prompting consistently outperforms English across all backbones, avoiding the semantic loss associated with translating specific administrative terminology ( p<0.05 , paired Student’s t-test over 3 independent runs).
Dimension
DGDB NEC retrieval
Primary Objective
Detecting implicit bias in official policy texts.
Core Domain
Legal and political administrative text ( high-stakes ).
Inventory
Dutch constitutional and equal-treatment law; antidiscrimination policy; accessibility and election communication guidance; online-discrimination, institutional-racism, emancipation, and inclusive-language resources; DGDB reasoning entries.
Selection Rationale
Grounds the review in legal norms while covering administrative, accessibility, and inclusive-language contexts that commonly shape implicit bias.
Appendix
Table 9: NEC retrieval construction for DGDB. MARS-Gov prioritizes normative sources and official writing guidance for administrative bias review.
Figure 6: ARG Hyperparameter Grid Search. (a) F1 Score and (b) Over-Governance Rate (OGR) across varying τhigh and juror-vote thresholds. The selected configuration ( τhigh=0.6,Nsoft_bias=2 ) is highlighted.