As LLM agents are increasingly deployed in multi-agent systems, they introduce risks of covert coordination that may evade standard forms of human oversight. While linear probes on model activations have shown promise for detecting deception in single-agent settings, collusion is inherently a multi-agent phenomenon, and the use of internal representations for detecting collusion between agents remains unexplored. We introduce NARCBench, a benchmark for evaluating collusion detection under environment distribution shift, and propose five probing techniques that aggregate per-agent deception scores to classify scenarios at the group level, evaluated across four open-weight models (Qwen3-32B, Llama-3.1-70B, DeepSeek-R1 32B, GPT-OSS-20B) and six probe architectures. We frame this as a distributed anomaly detection problem, identifying three collusion signatures that map onto distinct anomaly types and detection paradigms. Every model reaches 1.00 AUROC in-distribution; on our strongest model (Llama-3.1-70B), our five probing techniques achieve 0.73 to 0.93 AUROC when transferred zero-shot to structurally different multi-agent scenarios and 0.99 to 1.00 on a steganographic blackjack card-counting task, with detection performance scaling with model capability. We find that no single probing technique dominates across all collusion types, consistent with the framework's prediction that different anomaly types require different detection paradigms. This work takes a step toward multi-agent interpretability: extending white-box inspection from single models to multi-agent contexts, where detection requires aggregating signals across agents. These results suggest that model internals provide a complementary signal to text-level monitoring for detecting multi-agent collusion. Code and data available at https://github.com/aaronrose227/narcbench.
Figures & tables
Figure 1: Overview of our methodology. Activations are extracted from multi-agent interactions, projected onto a learned deception direction, and aggregated across agents to classify scenarios.
Figure 2: Distributed detection paradigms and how each processes multi-agent data to classify scenarios.
Figure 3: Per-test AUROC for each model on TRANSFER (solid; mean over six OOD families) and STEGO (hatched; twenty blackjack shoes). Means are taken across the model’s probe-layer range and 10 seeds.
Table 1: Zero-shot AUROC (mean ± std) of the five NARCBench Core detection methods on Verdict (Reflection phase).
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Sorted Concatenation
Act Alignment
Peak Suspicion
Deception Split
Asymmetry Probe
Core
1.00
1.00
1.00
1.00
1.00
Transfer (all)
0.82
0.48
0.84
0.88
0.87
CI/CD
0.53
0.28
0.32
0.73
0.48
Code Review
0.94
0.39
1.00
0.99
0.99
News Framing
0.71
0.71
0.94
0.96
0.94
Pump & Dump
0.99
0.32
0.97
0.83
0.96
Appendix
Table 2: Per-method AUROC across NARCBench tiers for Qwen3-32B. Bold = best per row. OOD Mean is computed across the available out-of-distribution scenarios ( Transfer families plus Stego ).
Sorted Concatenation
Act Alignment
Peak Suspicion
Deception Split
Asymmetry Probe
Core
1.00
0.97
1.00
1.00
1.00
Transfer (all)
0.89
0.73
0.90
0.92
0.93
CI/CD
0.96
0.84
0.89
0.89
0.92
Code Review
0.99
0.67
1.00
1.00
1.00
News Framing
0.93
0.83
0.96
0.97
0.99
Pump & Dump
0.95
0.93
0.98
0.99
0.99
Appendix
Table 3: Per-method AUROC across NARCBench tiers for Llama-3.1 70B. Bold = best per row.
Sorted Concatenation
Act Alignment
Peak Suspicion
Deception Split
Asymmetry Probe
Core
1.00
0.95
1.00
1.00
1.00
Transfer (all)
0.75
0.58
0.84
0.69
0.86
CI/CD
0.28
0.67
0.58
0.66
0.66
Code Review
0.87
0.91
0.91
0.93
0.93
News Framing
0.88
0.50
0.99
0.73
0.98
Pump & Dump
0.84
0.61
0.87
0.77
0.92
Appendix
Table 4: Per-method AUROC across NARCBench tiers for DeepSeek-R1 32B. Bold = best per row.
Sorted Concatenation
Act Alignment
Peak Suspicion
Deception Split
Asymmetry Probe
Core
1.00
0.91
1.00
1.00
1.00
Transfer (all)
0.62
0.72
0.56
0.65
0.64
CI/CD
0.76
0.98
0.77
0.85
0.83
Code Review
0.77
0.57
0.80
0.63
0.77
News Framing
−
−
−
−
−
Pump & Dump
−
−
−
−
−
Appendix
Table 5: Per-method AUROC across NARCBench tiers for GPT-OSS 20B. Bold = best per row. The News Framing, Pump & Dump and RAG Poisoning families produced no analysable collusion runs (every collusion attempt was refused or failed audit) and are marked “–”; Transfer (all) and OOD Mean are computed across the available families plus Stego .
Figure 5: OOD-mean AUROC ( Transfer + Stego ) plotted against model parameter count. Detection performance scales approximately monotonically with model capability across the four open-weight models we tested.
Probe
Qwen3-32B
Llama-3.1-70B
DeepSeek-R1-32B
GPT-OSS-20B
Logistic Regression
0.81
0.89
0.77
0.62
Multi-Layer Logistic Regression
0.72
0.87
0.78
0.60
Difference of Means
0.68
0.86
0.67
0.68
Multi-Headed Attention
0.60
0.73
0.60
0.64
Multi-Layer Perceptron
0.56
0.77
0.68
0.57
Linear Artificial Tomography
0.52
0.57
0.54
0.63
Appendix
Table 6: OOD-mean AUROC per probe architecture per model (averaged across the six Transfer families and Stego ). Bold = highest per model.
Transfer
Stego
Method
∣Gcol∣=2
∣Gcol∣=3
∣Gcol∣=4
∣Gcol∣=2
∣Gcol∣=3
∣Gcol∣=4
Sorted Concat
0.635
0.711
0.737
0.998
1.000
0.999
Activation Alignment
0.446
0.443
0.418
1.000
1.000
1.000
Peak Suspicion
0.506
0.627
0.623
1.000
1.000
1.000
Deception Split
0.575
0.682
0.560
1.000
1.000
1.000
Asymmetry Probe
0.568
0.688
0.589
1.000
1.000
1.000
Appendix
Table 7: Mean AUROC for each colluding group size ∣Gcol∣ , on the Transfer and Stego evaluation sets.
Method
Eval-Aware OOD Mean AUROC
Non-Eval-Aware OOD Mean AUROC
Sorted Concat
0.779
0.846
Activation Alignment
0.660
0.552
Peak Suspicion
0.833
0.863
Deception Split
0.819
0.893
Asymmetry Probe
0.834
0.886
Appendix
Table 8: Mean AUROC on out-of-distribution (OOD) evaluation sets, under eval-aware and non-eval-aware generation.
Group
n
Mean ρ
Colluder, Big Player (seat 1)
20
+0.416
Colluder, Counter (seat 0)
20
−0.025
Honest, collusion runs
40
−0.020
Honest, control runs
80
−0.037
Appendix
Table 10: Mean lag-1 bet/count correlation by group across all 40 shoes.
Figure 7: Token-level analysis in NARCBench-Stego . (a) Real result: the Big Player (colluder) projects higher onto the deception direction than honest agents reading the same Counter message ( +0.29 vs −1.41 across n=39 matched examples). (b) Schematic, not measured: illustrative per-token heatmap showing how the effect in (a) would look distributed across the tokens of the Counter’s message. Colours match the paper’s convention: red for colluder roles, blue for honest.
Method
Exploit-Share
Sorted Concat
0.649±0.010
Activation Alignment
0.709±0.001
Peak Suspicion
0.554±0.006
Deception Split
0.737±0.012
Asymmetry Probe
0.565±0.012
Appendix
Table 12: Zero-shot AUROC of the five probe methods, trained on Core (qwen3_32b) cross-model on Exploit-Share ( → gpt-oss-120b)
Figure 8: Per-method AUROC on Llama-3.1-70B across Held-Out , Transfer , and Stego . Means over probe-layer range and ten seeds; error bars show layer standard deviation.
Model
Transfer
Stego
Llama-3.1 70B
0.87
1.00
Qwen3 32B
0.78
1.00
DeepSeek-R1 32B
0.75
0.95
GPT-OSS 20B
0.68
0.86
Appendix
Table 13: Mean AUROC per model, averaged across the five probing techniques, seeds, layers (and Transfer families).