cs.AIAug 4, 2026

Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems

Authors: Sebastián Andrés Cajas OrdóñezAgastya MunnangiAldo MarzulloFelipe Ocampo OsorioQuang BuiMohammad ShahinArmaan GrewalEmmanuel Paul Kwesiga+8 more

Organizations: MIT Critical Data, Massachusetts Institute of Technology, Cambridge, Massachusetts, United States · Georgia State University, Atlanta, Georgia, United States · Politecnico di Milano, Milan, Italy · American International School Vienna, Vienna, Austria · School of Public Health, Boston University, Boston, Massachusetts, United States · Northwestern University, Evanston, Illinois, United States · Technische Hochschule Lübeck, Lübeck, Germany · Substrate Labs · School of Nursing, University of British Columbia, Vancouver, British Columbia, Canada · Dartmouth College, Hanover, New Hampshire, United States · Department of Computer Science, University of Maryland, College Park, Maryland, United States · McGill University, Montreal, Quebec, Canada · Department of Molecular Biosciences, University of Texas at Austin, Austin, Texas, United States · King’s College London, London, United Kingdom · Division of Pulmonary, Critical Care and Sleep Medicine, Beth Israel Deaconess Medical Center, Boston, Massachusetts, United States

Abstract

Clinical decision support is moving toward committees of language-model agents deliberating on a shared workspace. We ask whether such committees can be gamed by shortcuts, cues a benchmark rewards but a clinician would ignore. Across seven cohorts on six public datasets spanning text (MedQA-USMLE, MedMCQA, MIMIC-CXR reports), imaging (NIH ChestX-ray14, MIMIC-CXR-JPG, CheXpert) and tabular ICU records (SUPPORT2), Gemini committees resist these cues in isolation (flip 5-16%), yet a socially plausible shortcut spreads: when two peers assert the same wrong answer, the holdout under test adopts it in 38% of cases, as does a false "pre-screen" system flag, on both capability tiers. Of three oversight agents, a gate cannot separate adoption from honest agreement (false-positive rate 100%); a same-lineage judge reading only the transcript flags adoption on text (precision 100%, recall 93%) but collapses onto the gate in imaging; a referee that privately re-queries the holdout transfers to imaging (77-88% precision, 13-21% false-positive rate). Tripling a cue's visual salience does not move contagion, whereas a second peer voice raises it by half again. Gaming a hidden rubric is near-silent: only 1/10 text and 1/134 imaging drifters name the rubric they moved toward. What games a committee is social plausibility, and only a referee independent of self-report catches it. Code: https://github.com/criticaldata/benchmaxxing

Explore similar work

CardsList
  1. World Feedback for Clinical Agents: Diagnosing RL in FHIR Environments

    Jul 1, 2026Ananya Mantravadi, Harshit Rajgarhia, Prasanna Desikan +1Opti-Agent-BenchFeedback Learning