Semantic Cooperative Games for Contribution Attribution in LLM-Based Multi-Agent Systems
Organizations: New York University · University of California, Davis
Abstract
Contribution attribution has become a central problem in LLM-based multi-agent systems, where final outputs are produced through multiple agents, message exchanges, and ordered workflow dependencies. Existing attribution methods often rely on counterfactual valuation, such as removing agents or comparing score changes across altered agent subsets. In language-mediated workflows, these methods require repeated model calls, introduce high variance, and do not explicitly capture the intermediate semantic states through which agents produce, preserve, and transform task-relevant information. We propose Semantic Cooperative Games (SCG), a framework that represents a realized language flow as a semantic generation hypergraph and induces an agent-level semantic value function on this structure. We define the Semantic Shapley Value (SSV) to allocate contribution over semantic support logic, and introduce SLIC, a single-trajectory algorithm that constructs the semantic hypergraph, recovers minimal semantic supports, applies Boolean absorption, and computes SSV without rerunning agent subsets. We prove that SSV reduces to the classical Shapley value under standard set-based, fully observable, and no-order-dependence conditions. On a medical benchmark satisfying these conditions, SLIC reduces computation cost by 93.3% while remaining highly consistent with a Monte Carlo Shapley baseline. In more general multi-role workflows, SSV aligns with perturbation-induced score-drop profiles and exposes cases where semantic contribution and failure impact diverge. Overall, SLIC provides a fast, counterfactual-free, and interpretable attribution method for complex LLM-based multi-agent systems.
Figures & tables
| Method | Kendall | L1 Error | API Calls (per case) | Cost Reduction |
|---|---|---|---|---|
| Monte Carlo (GT) | 1.000 | 0.00 | (avg 95.9) | Baseline |
| Leave-One-Out (LOO) | 0.575 | 15.08 | (avg 54.8) | 42.9% |
| Holistic LLM Judge | 0.479 | 12.90 | (avg 13.7) | 85.7% |
| SLIC (Ours) | 0.814 | 6.61 | (avg 13.7) | 85.7% |
| Setting | Weak | Medium | Strong | LOO / C3-style |
|---|---|---|---|---|
| Meeting decision summarization | 0.718 | 1.000 | 1.000 | 0.975 |
| Table-text numerical reasoning | 0.667 | 0.900 | 0.900 | 0.900 |
| Rule application | 0.500 | 0.700 | 0.700 | 0.700 |
| HealthBench-style medical safety | 0.564 | 0.872 | 0.872 | 0.667 |
| Mean | 0.612 | 0.868 | 0.868 | 0.810 |
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
| Relation | Planned | Valid | Confirmed | Rate [95% CI] | Type match | |||
|---|---|---|---|---|---|---|---|---|
| Transformed | 120 | 120 | 110 | .917 [.867, .958] | .950 | .028 | .956 | 118/120 |
| Preserved | 40 | 40 | 39 | .975 [.925, 1.000] | .983 | .000 | .975 | 40/40 |
| New | 160 | 159 | 154 | .969 [.937, .994] | .969 | .969 | – | 159/159 |
| Task family | Support conf. | New conf. | |||||
|---|---|---|---|---|---|---|---|
| HealthBench-style | 40 | 36/40 | .917 | .008 | .925 | 39 | 39/39 |
| LegalBench-style | 40 | 37/40 | .958 | .025 | .967 | 40 | 37/40 |
| QMSum-style | 40 | 38/40 | .958 | .000 | .950 | 40 | 38/40 |
| TAT-QA-style | 40 | 38/40 | 1.000 | .050 | 1.000 | 40 | 40/40 |
| Overall | 160 | 149/160 | .958 | .021 | .960 | 159 | 154/159 |
| Analyzer | SSV rank | SSV | Raw-link Jaccard | Link–agent Jaccard |
|---|---|---|---|---|
| Gemini 2.5 Flash-Lite | ||||
| Gemini 2.5 Flash | ||||
| Llama 4 Scout | ||||
| Llama 4 Maverick |
| Method | Kendall | L1 Error | API Calls (per case) | Cost Reduction |
|---|---|---|---|---|
| Monte Carlo (GT) | 1.000 | 0.00 | Baseline | |
| LOO (Duplicated) | 0.000 | 21.18 | 42.9% | |
| Holistic LLM Judge | 0.479 | 12.90 | 85.7% | |
| SLIC (Ours) | 0.814 | 6.61 | 85.7% |
| Case ID | Total Rubrics | Grand Coalition Count | Any-Subset Count |
| 0 | 19 | 1 | 3 |
| 1 | 10 | 2 | 2 |
| 3 | 20 | 0 | 3 |
| 4 | 13 | 0 | 2 |
| 5 | 12 | 1 | 4 |
| 6 | 14 | 1 | 3 |
| Method | API passes / case | API Calls | Kendall | L1 Error |
|---|---|---|---|---|
| MC (GT) | 15 | 207.0 | 1.000 | 0.000 |
| LOO | 5 | 69.0 | 0.443 | 14.433 |
| LOO redundant endpoint | 5 | 69.0 | 0.000 | 17.000 |
| Holistic | 1 | 13.8 | 0.499 | 7.467 |
| SLIC | 1 | 13.8 | 0.665 | 5.700 |
| Setting | Weak | Medium | Strong | LOO / C3-style | |
|---|---|---|---|---|---|
| Meeting decision summarization | 50 | 0.718 | 1.000 | 1.000 | 0.975 |
| Table-text numerical reasoning | 50 | 0.667 | 0.900 | 0.900 | 0.900 |
| Rule application | 50 | 0.500 | 0.700 | 0.700 | 0.700 |
| HealthBench-style medical safety | 50 | 0.564 | 0.872 | 0.872 | 0.667 |
| Mean | 200 | 0.612 | 0.868 | 0.868 | 0.810 |
| Series | A | B | C | D | E |
|---|---|---|---|---|---|
| SSV | 0.345 | 0.376 | 0.279 | 0.000 | 0.000 |
| Weak | 0.510 | 0.093 | 0.004 | 0.013 | 0.000 |
| Medium | 0.348 | 0.387 | 0.265 | 0.000 | 0.000 |
| Strong | 0.350 | 0.385 | 0.265 | 0.000 | 0.000 |
| LOO | 0.354 | 0.382 | 0.262 | 0.002 | 0.000 |
| Series | A | B | C | D | E |
|---|---|---|---|---|---|
| SSV | 0.539 | 0.342 | 0.117 | 0.001 | 0.000 |
| Weak | 0.048 | 0.578 | 0.134 | 0.000 | 0.000 |
| Medium | 0.536 | 0.137 | 0.306 | 0.020 | 0.000 |
| Strong | 0.526 | 0.138 | 0.314 | 0.022 | 0.000 |
| LOO | 0.893 | 0.082 | 0.006 | 0.019 | 0.000 |
| Series | A | B | C | D | E |
|---|---|---|---|---|---|
| SSV | 0.299 | 0.200 | 0.384 | 0.117 | 0.000 |
| Weak | 0.087 | 0.470 | 0.215 | 0.089 | 0.000 |
| Medium | 0.251 | 0.174 | 0.305 | 0.269 | 0.000 |
| Strong | 0.253 | 0.182 | 0.300 | 0.265 | 0.000 |
| LOO | 0.269 | 0.145 | 0.298 | 0.288 | 0.000 |
| Series | A | B | C | D | E |
|---|---|---|---|---|---|
| SSV | 0.392 | 0.000 | 0.384 | 0.224 | 0.000 |
| Weak | 0.017 | 0.007 | 0.414 | 0.562 | 0.000 |
| Medium | 0.329 | 0.008 | 0.346 | 0.316 | 0.000 |
| Strong | 0.334 | 0.005 | 0.346 | 0.316 | 0.000 |
| LOO | 0.152 | 0.012 | 0.052 | 0.783 | 0.000 |
| Setting | Flow | Agent responsibilities |
|---|---|---|
| Meeting decision summarization | A/B/C D E | A: Decision and scope B: Constraints and blockers C: Action items and owners D: Review, backup options, advice E: Transparent pass-through |
| Table-text numerical reasoning | A/B C D E | A: Targets and values B: Computation rules C: A/B-based computation D: Explanation and final answer E: Transparent pass-through |
| Rule application | A/B/C D E | A: Claim focus B: Main rule C: Conditions and exceptions D: Rule application and final judgment E: Transparent pass-through |
| HealthBench medical safety | A/B C D E | A: Raw case facts B: Formatting and remote-answer boundary C: Risk assessment D: Next-step action and safety boundary E: Transparent pass-through |
| Experiment group | Function | Model |
|---|---|---|
| Four-scenario radar | Agent generation | models/gemini-2.5-flash |
| Four-scenario radar | Rubric judge | models/gemini-2.5-flash |
| Four-scenario radar | Semantic analyzer / link extraction | models/gemini-2.5-flash |
| Four-scenario radar | Attack / rerun downstream generation | models/gemini-2.5-flash |
| HealthBench Shapley-alignment | Agent generation / merge | models/gemini-2.5-flash |
| HealthBench Shapley-alignment | Rubric scoring / rejudge | models/gemini-2.5-pro |