MASCRDM: Multi-Agent System for Compliance Risk Detection and Mitigation in Training Process of Large Language Models
Organizations: Department of Computer Science and Technology, Tsinghua University · Academy of Artificial Intelligence and Advanced Technology, Xi’an Jiaotong-Liverpool University · Department of Computer Sciences, University of Wisconsin–Madison · Law School, University of Chinese Academy of Social Sciences
Abstract
Large Language Models (LLMs) have been applied in various fields. However, ensuring compliance and safety of LLMs, such as avoiding discrimination and bias, still remains a challenge. Current efforts mainly focus on detecting and filtering inputs and outputs of the trained models, rather than studying the intrinsic architecture of the models in real-time. To tackle this challenge, we analyze the LLMs training process and discover two critical issues: 1) Most of the existing methods are predominantly static in their approach to detection and filtering, achieving only localized optimizations without systematically enhancing the compliance of LLMs. 2) Another issue with existing approaches is the lack of real-time risk detection and mitigation across the full training process, which leads to limited flexibility. Motivated by these, we propose MASCRDM (Multi-Agent System for Compliance Risk Detection and Mitigation) during the LLM training process. Firstly, we develop a set of compliance rules based on existing Artificial Intelligence (AI) laws and a compliance-specific LLM with the instruction of compliance law experts. Then, we deconstruct LLMs into several components and identify key nodes based on the compliance knowledge graph. During LLMs training, we implement our multiple agents in the whole process, giving compliance risk alerts and suggestions for LLM developers. Experiments on discrimination and bias benchmark demonstrate that our multi-agent system can effectively improve the compliance while maintaining reasonable semantic performance. The results indicate that our method provides an executable path for mitigating compliance risk from within the LLMs systematically.
Figures & tables
| Method | BBQ | CrowS-Pairs | ||||
|---|---|---|---|---|---|---|
| Acc. | A.Amb | A.Dis | B.Amb | B.Dis | SS | |
| ( ) | ( ) | ( ) | ( ) | ( ) | ( ) | |
| Qwen3-8B-Base | 61.42 | 46.17 | 76.67 | 3.99 | 1.02 | 64.46 |
| KLAAD | 62.09 | 45.65 | 78.52 | 5.94 | 0.83 | 69.56 |
| Fairness mediator | 59.46 | 45.01 | 73.90 | 3.41 | 1.43 | 62.40 |
| Bias Unlearning | 63.01 | 52.13 | 73.89 | 3.37 | 1.11 | 61.74 |
| Type | Method | Sentiment | VAD | BE5 | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| V | A | D | Joy | Anger | Sadness | Fear | Disgust | |||
| Profession (Engineering Branches) | Qwen3-8B-Base | +0.12 | +0.32 | -0.13 | +0.27 | 0.22 | 0.14 | 0.14 | 0.16 | 0.14 |
| KLAAD | +0.21 | +0.33 | -0.13 | +0.23 | 0.24 | 0.15 | 0.15 | 0.17 | 0.15 | |
| Fairness mediator | +0.13 | +0.33 | -0.14 | +0.28 | 0.22 | 0.14 | 0.14 | 0.16 | 0.14 | |
| Bias Unlearning | +0.13 | +0.33 | -0.14 | +0.27 | 0.22 | 0.14 | 0.14 | 0.15 | 0.14 | |
| MASCRDM | +0.12 | +0.32 | -0.13 | +0.27 | 0.22 | 0.14 | 0.14 | 0.15 | 0.14 | |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Parameter | Value |
|---|---|
| Training data | Alpaca + Stereo |
| Visible GPUs | 4 A800 |
| Maximum sequence length | 1024 |
| Micro-batch size | 4 |
| Gradient accumulation steps | 8 |
| Effective batch size per update | 32 |
| Parameter | Value |
|---|---|
| Audit frequency | Every global step |
| Probe set | 36 S/C/N triplets |
| Embedding trigger | / |
| Embedding correction | |
| Attention threshold | P90 |
| Attention strength |
| Method | BBQ | |||||
|---|---|---|---|---|---|---|
| Acc. | A.Amb | A.Dis | B.Amb | B.Dis | ||
| ( ) | ( ) | ( ) | ( ) | ( ) | ||
| Qwen3-8B-Base | 60.94 | 47.54 | 74.34 | 3.59 | 1.13 | |
| Qwen3-8B-Base+Data Audit | 63.41 | 50.36 | 76.45 | 3.76 | 0.92 | |
| +Embedding | 63.43 | 50.57 | 76.29 | 3.73 | 1.10 | |
| +Attention | 62.88 | 52.53 | 73.22 | 3.37 | 1.00 | |
| Type | Method | Sentiment | VAD | BE5 | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| V | A | D | Joy | Anger | Sadness | Fear | Disgust | |||
| Gender (Male) | Qwen3-8B-Base | +0.30 | +0.43 | -0.06 | +0.30 | 0.24 | 0.14 | 0.14 | 0.16 | 0.13 |
| KLAAD | +0.72 | +0.60 | +0.03 | +0.44 | 0.29 | 0.14 | 0.14 | 0.16 | 0.13 | |
| Fairness mediator | +0.21 | +0.42 | -0.05 | +0.28 | 0.24 | 0.14 | 0.14 | 0.16 | 0.14 | |
| Bias Unlearning | +0.28 | +0.43 | -0.06 | +0.28 | 0.24 | 0.14 | 0.14 | 0.15 | 0.13 | |
| MASCRDM | +0.29 | +0.44 | -0.07 | +0.29 | 0.24 | 0.14 | 0.14 | 0.15 | 0.13 | |
| Type | Method | Sentiment | VAD | BE5 | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| V | A | D | Joy | Anger | Sadness | Fear | Disgust | |||
| Political Ideology (Nationalism) | Qwen3-8B-Base | +0.15 | +0.28 | +0.01 | +0.41 | 0.21 | 0.15 | 0.15 | 0.17 | 0.15 |
| KLAAD | +0.30 | +0.32 | -0.03 | +0.37 | 0.24 | 0.16 | 0.16 | 0.17 | 0.15 | |
| Fairness mediator | +0.13 | +0.31 | -0.01 | +0.41 | 0.22 | 0.15 | 0.15 | 0.17 | 0.15 | |
| Bias Unlearning | +0.12 | +0.30 | -0.02 | +0.40 | 0.21 | 0.15 | 0.15 | 0.16 | 0.15 | |
| MASCRDM | +0.13 | +0.29 | -0.02 | +0.41 | 0.21 | 0.15 | 0.15 | 0.16 | 0.15 | |
| Type | Method | Sentiment | VAD | BE5 | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| V | A | D | Joy | Anger | Sadness | Fear | Disgust | |||
| Profession (Film And Television Occupations) | Qwen3-8B-Base | +0.23 | +0.42 | -0.11 | +0.28 | 0.24 | 0.14 | 0.14 | 0.15 | 0.14 |
| KLAAD | +0.49 | +0.47 | -0.16 | +0.31 | 0.28 | 0.14 | 0.14 | 0.16 | 0.14 | |
| Fairness mediator | +0.20 | +0.41 | -0.11 | +0.27 | 0.23 | 0.14 | 0.14 | 0.15 | 0.14 | |
| Bias Unlearning | +0.19 | +0.42 | -0.11 | +0.27 | 0.24 | 0.14 | 0.14 | 0.15 | 0.14 | |
| MASCRDM | +0.19 | +0.42 | -0.12 | +0.26 | 0.23 | 0.14 | 0.14 | 0.16 | 0.14 | |
| Type | Method | Sentiment | VAD | BE5 | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| V | A | D | Joy | Anger | Sadness | Fear | Disgust | |||
| Profession (Theatre Personnel) | Qwen3-8B-Base | +0.24 | +0.42 | -0.10 | +0.30 | 0.23 | 0.14 | 0.14 | 0.15 | 0.14 |
| KLAAD | +0.40 | +0.47 | -0.12 | +0.29 | 0.27 | 0.14 | 0.14 | 0.16 | 0.14 | |
| Fairness mediator | +0.23 | +0.43 | -0.09 | +0.30 | 0.24 | 0.14 | 0.14 | 0.15 | 0.14 | |
| Bias Unlearning | +0.23 | +0.42 | -0.10 | +0.28 | 0.23 | 0.14 | 0.14 | 0.15 | 0.14 | |
| MASCRDM | +0.22 | +0.42 | -0.09 | +0.30 | 0.23 | 0.14 | 0.14 | 0.15 | 0.14 | |
| Type | Method | Sentiment | VAD | BE5 | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| V | A | D | Joy | Anger | Sadness | Fear | Disgust | |||
| Religious Ideology (Islam) | Qwen3-8B-Base | +0.11 | +0.31 | -0.15 | +0.31 | 0.22 | 0.16 | 0.16 | 0.17 | 0.15 |
| KLAAD | +0.06 | +0.19 | -0.02 | +0.30 | 0.26 | 0.19 | 0.19 | 0.21 | 0.18 | |
| Fairness mediator | +0.11 | +0.36 | -0.11 | +0.38 | 0.23 | 0.16 | 0.16 | 0.18 | 0.16 | |
| Bias Unlearning | +0.12 | +0.32 | -0.12 | +0.36 | 0.22 | 0.16 | 0.16 | 0.17 | 0.15 | |
| MASCRDM | +0.08 | +0.33 | -0.17 | +0.31 | 0.21 | 0.16 | 0.16 | 0.17 | 0.16 | |