Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety
Organizations: Harvard University
Abstract
Safety benchmarks usually test "bare" models that receive prompts and output responses, but real-world deployments "wrap" those models in complex scaffolds. How much do these scaffolds affect model safety as measured by benchmarks? We test six leading models on four pre-registered safety benchmarks with a direct API and three scaffolds: ReAct, multi-agent, and map-reduce. We conducted 60,112 scored evaluations. On average, how safety is measured matters more than scaffolding does: we find that using a multiple choice vs. open-ended format for otherwise-identical benchmark items changes measured safety by about 5-20 percentage points (pp). The two formats are scored with different methods (answer extraction and an LLM judge), so the gap is due to measurement rather than differences in latent safety. Using a heuristic to classify model refusals would have led to different findings in four of five cases. Benchmark choice explains 15.1% of the variation in outcomes; scaffold architecture explains 0.5%, about 33x less. We find that map-reduce scaffolds, a form of structure-destroying delegation that strips answer options by decomposing prompts, reduce pooled measured safety by 7.3 pp (95% CI: 6.4 to 8.1). The pooled effects for ReAct and multi-agent scaffolds are within our pre-registered +/-2 pp margin of equivalence. However, there are large differences across models for specific benchmarks and scaffolds that are hidden by pooled estimates: for example, on the same sycophancy benchmark items, Opus 4.6 has 16.8 pp lower measured safety with a map-reduce scaffold, while Llama 4 has 18.8 pp higher measured safety. Composite reliability is G = 0.251 (95% CI: [0.000, 0.879]). This wide confidence interval, which spans "of little use" to "very good", does not support using a single composite measure of model safety as the basis for go/no-go decisions about model deployment.
Figures & tables
| Study | Cases models configs | records | Pre-reg status | Items / scoring |
|---|---|---|---|---|
| Primary scaffold eval | (not fully crossed) | Pre-registered (CJW92; sycophancy added by addendum) | BBQ 800; TQA 817; XSTest 500; Sycophancy 500. Deterministic MC + LLM-judge for XSTest. | |
| AI factual-recall control | Pre-registered (CJW92) | 500 AI/ML factual-knowledge items from Anthropic MWE (originally mis-labelled as “sycophancy”; repurposed as a non-safety control). Deterministic MC. | ||
| Phase 2 mechanistic | per benchmark | Pre-registered (WA9Y7) | 300 BBQ + 300 TQA items, all also in the primary dataset; 4 invocation conditions. | |
| Format dependence | Pre-data exploratory; design frozen pre-collection | Paired MC/OE on BBQ 60, TQA 30, syc 70, AIFR 10, MMLU 50. Direct + full-question aggregation. |
| ID | Element | Pre-registered | Implemented |
|---|---|---|---|
| D-001 | Prompt placement | Benchmark instructions via system prompt (§PR 9.1 to 9.2) | Resolved : implemented as registered (benchmark instructions as system prompt, item as user message; §3.3) |
| D-002 | XSTest scoring | LLM judge (Gemini 3 Flash primary, §PR 7.3) | Resolved : LLM judge (Gemini 3 Flash primary, Opus 4.6 validation) now used as pre-registered |
| D-003 | Multi-agent rounds | Max 1 revision round (§PR 4.1) | Resolved : implemented as registered (one revision round) |
| D-004 | ReAct tools | 3 tools: calculator, text_search, scratchpad (§PR 4.1) | Resolved : implemented as registered (local tools; no external retrieval) |
| D-005 | Spec curve scope | 1,000 to 2,000 specs, 500 perms (§PR 5.4) | 18 primary specs (3 DOF), 384 exploratory specs (9 DOF) |
| D-006 | Primary estimator | GLMM with case random intercept (§PR 5.1) | Case-cluster-robust logistic regression; saved GLMM fit diagnostics unavailable; pre-registered LRT comparisons for H2 and H3 accordingly conducted as cluster-robust Wald tests |
| Comparison | OR [95% CI] | (raw) | (Holm) | RD [95% CI] | NNH |
|---|---|---|---|---|---|
| ReAct vs. Direct ∗ | 0.96 [0.92, 0.99] | 0.009 | 0.017 | [ , ] | 133 |
| Multi-Agent vs. Direct ∗ | 0.95 [0.91, 0.99] | 0.010 | 0.017 | [ , ] | 119 |
| Map-Reduce vs. Direct ∗∗∗ | 0.64 [0.61, 0.68] | 0.001 | 0.001 | [ , ] | 14 |
| ; (Holm-corrected). | |||||
| ReAct | Multi-agent | Map-reduce | ||||
| Model | Direct | RD (pp) | Direct | RD (pp) | Direct | RD (pp) |
| DeepSeek V3.2 | 70.2% | 70.2% | 70.2% | |||
| GPT-5.2 | 73.4% | 73.4% | 73.4% | |||
| Llama 4 | 67.0% | 67.0% | 67.0% | |||
| Mistral Large 2 | 72.9% | 72.9% | 72.9% | |||
| Opus 4.6 | 85.2% | 85.2% | 85.2% | |||
| Benchmark | MC (pooled) | OE (pooled) | Gap (pp) | Direction |
|---|---|---|---|---|
| Sycophancy | 33.7% | 53.3% | Higher OE score | |
| BBQ | 94.3% | 99.2% | Higher OE score | |
| TruthfulQA | 79.3% | 85.0% | Higher OE score | |
| AI Factual Recall | 77.0% | 76.0% | Negative-control contrast | |
| MMLU (capability) | 85.4% | 76.2% | Higher MC score |
| Model | MC | OE | Gap (pp) |
|---|---|---|---|
| DeepSeek V3.2 | 96.7% | 100.0% | |
| GPT-5.2 | 90.0% | 99.2% | |
| Llama 4 Maverick | 94.2% | 97.5% | |
| Mistral Large 2 | 95.0% | 99.2% | |
| Opus 4.6 | 95.8% | 100.0% |
| Model | MC | OE | Gap (pp) |
|---|---|---|---|
| DeepSeek V3.2 | 32.1% | 48.6% | |
| GPT-5.2 | 57.1% | 60.0% | |
| Llama 4 Maverick | 15.0% | 29.3% | |
| Mistral Large 2 | 22.1% | 53.6% | |
| Opus 4.6 | 42.1% | 75.0% |
| Safe of 200 | |||||
|---|---|---|---|---|---|
| Model | Direct | MR | OP | OP MR (pp) | Recovery (%) |
| Opus 4.6 | 189 | 157 | 186 | [9.5, 20.0] | 90.6 [75.0, 103.4] |
| DeepSeek V3.2 | 162 | 101 | 141 | [13.0, 27.0] | 65.6 [47.8, 82.4] |
| Mistral Large 2 | 169 | 134 | 156 | [5.0, 17.0] | 62.9 [34.9, 90.0] |
| GPT-5.2 | 176 | 146 | 158 | [0.5, 11.5] | 40.0 [3.6, 66.7] |
| Llama 4 Maverick | 158 | 143 | 143 | [ , 6.0] | 0.0 (CI unstable) |
| Sub-call type | Option-marker hits | System-keyword hits | |
|---|---|---|---|
| Decompose | 150 | 75 | 3 |
| Map chunk 0 | 150 | 0 | 150 |
| Map chunk 1 | 148 | 3 | 148 |
| Map chunk 2 | 125 | 5 | 125 |
| Reduce | 150 | 75 | 150 |
| Total | 723 | 158 | 576 |
| Model | Direct | CoT | CoT–Direct | MR–Direct | Interpretation |
|---|---|---|---|---|---|
| Gemini 3 Pro | 84.4% | 89.9% | pp | pp | CoT benefits modestly |
| Opus 4.6 | 88.8% | 91.4% | pp | pp | CoT benefits modestly |
| Llama 4 | 75.4% | 76.5% | pp | pp | CoT benefits slightly |
| GPT-5.2 | 83.5% | 83.9% | pp | pp | CoT neutral |
| DeepSeek V3.2 | 76.4% | 76.0% | pp | pp | CoT neutral |
| Model | PT | Min | Mod | Agg | (Agg–PT) |
|---|---|---|---|---|---|
| Opus 4.6 | 96.3% | 95.0% | 93.7% | 94.0% | pp |
| Gemini 3 Pro | 93.7% | 94.0% | 94.7% | 89.7% | pp |
| GPT-5.2 | 92.7% | 89.3% | 85.7% | 79.7% | pp |
| DeepSeek V3.2 | 91.3% | 91.0% | 83.3% | 79.0% | pp |
| Llama 4 Mav | 92.0% | 93.0% | 92.0% | 87.7% | pp |
| Mistral Large 2 | 88.7% | 89.3% | 78.7% | 66.7% | pp |
| Model | PT | Min | Mod | Agg | (Agg–PT) |
|---|---|---|---|---|---|
| Opus 4.6 | 98.3% | 97.7% | 98.3% | 98.7% | pp |
| Gemini 3 Pro | 92.7% | 92.3% | 94.3% | 97.0% | pp |
| GPT-5.2 | 90.7% | 86.3% | 88.7% | 93.7% | pp |
| DeepSeek V3.2 | 79.7% | 79.7% | 82.3% | 91.7% | pp |
| Llama 4 Mav | 69.0% | 79.3% | 82.7% | 87.7% | pp |
| Mistral Large 2 | 83.7% | 79.7% | 84.3% | 90.0% | pp |
| Safety Rate | RD vs. Direct (pp) | |||||
|---|---|---|---|---|---|---|
| Model | Config | PF% | ITT | PP | ITT | PP |
| Opus 4.6 | Direct | 0.3% | 84.5% | 84.7% | — | — |
| ReAct | 0.2% | 83.3% | 83.5% | 1.2 | 1.3 | |
| Multi-agent | 0.3% | 83.5% | 83.7% | 1.0 | 1.0 | |
| Map-reduce | 5.7% | 69.8% | 74.0% | 14.7 | 10.7 | |
| GPT-5.2 | Direct | 1.1% | 77.4% | 78.3% | — | — |
| Model | Direct | ReAct | Multi-Agent | Map-Reduce | |
|---|---|---|---|---|---|
| GPT-5.2 | 42.5 | 41.6 | 45.5 | 45.8 | |
| Opus 4.6 | 49.0 | 47.2 | 50.2 | 32.2 | |
| DeepSeek V3.2 | 34.6 | 41.4 | 37.0 | 30.4 | |
| Mistral Large 2 | 32.2 | 30.0 | 31.8 | 33.4 | |
| Llama 4 | 11.0 | 10.8 | 15.6 | 29.8 | |
| Pooled | 33.8 | 34.1 | 36.0 | 34.2 |
| Model | Direct | MR NoLeak | MR Leak | Gap | %Leak |
| Opus 4.6 | 49.0 | 46.0 | 18.7 | 50.4% | |
| GPT-5.2 | 42.5 | 47.0 | 42.3 | 26.9% | |
| DeepSeek V3.2 | 34.6 | 34.8 | 27.2 | 58.0% | |
| Mistral Large 2 | 32.2 | 35.0 | 29.5 | 29.8% | |
| Llama 4 | 11.0 | 35.8 | 12.4 | 25.8% | |
| Pooled | 33.8 | 39.7 | 25.4 | 38.3% |
| Finding | Heuristic | Judge | Verdict |
|---|---|---|---|
| Over-refusal (MR Direct) | pp | pp | Consistent |
| Agreement collapse (3-way exact) | pp | pp | Reversed |
| DeepSeek boundary softening | pp | pp | Reversed |
| Sub-call information leakage rate | 39.3% | 4.4% | Mostly scoring artifact |
| Safety architecture (DSS) distinction | 3 distinct types | Clear-cut: distinct; boundary: concentrated | Refined |
| Threat Channel | Mitigation | Residual Risk |
|---|---|---|
| Opus test scores contaminate results | Opus-excluded sensitivity analysis ( ; Table 17 ) | Low |
| Scoring rubrics favour Opus-like outputs | Rubrics derived from published benchmark criteria; identical prompts across all models | Low; auditable |
| Pipeline micro-decisions (prompt formatting, retry logic, answer-extraction) iteratively developed using Opus may advantage Opus | All models share a single LiteLLM code path with provider-specific calls confined to API-parameter handling (Appendix J ); no model-specific branching in prompt, scoring, or extraction logic; code publicly released | Medium; requires independent replication |
| Opus’s robustness is an artifact of prompts Opus designed | Key finding (refusal under invocation) is binary behavioural outcome; adversarial prompts from published benchmarks | Low–Medium |
| Test | Metric | Full (6 models) | Excl. Opus (5 models) | Change? |
| H1a (ReAct) | OR | 0.96 | 0.97 | Slightly closer to 1 |
| RD (pp) | Less negative | |||
| 0.017 | 0.16 | Sig. NS | ||
| H1b (Multi-agent) | OR | 0.95 | 0.95 | No |
| RD (pp) | Less negative | |||
| 0.017 | 0.086 | Sig. NS |
| Recommendation | Evidence Basis | Practice Gap | Cost |
|---|---|---|---|
| Format-paired reporting | about 5 to 20 pp format gaps (Sec. 5.1 ) | No benchmark requires dual format | Low |
| Structure-destroying scaffold test | NNH = 14 under map-reduce (Sec. 4.2 ) | Proxy safety benchmarks not scaffold-tested | Medium |
| Propagation verification | Option-marker loss at map workers (Sec. 5.2 ) | No standard propagation audit | Low |
| NNH operational reporting | Enterprise risk communication (Sec. 7.2 ) | Safety scores lack operational interpretation | Low |
| System prompt governance | Prompt competition can suppress scaffolds (Sec. 18 ) | No API mechanism for safety-priority prompts | Low |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Dimension | Map-Reduce | RLM |
| MC choices in sub-calls | 0% | 77.9–87.9% (context var.) |
| System prompt propagation | Yes | 0% |
| Content transformation | Abstract sub-questions | Full context via variable ref. |
| Sub-call count | 3–5 (fixed pipeline) | 3–12 (model-decided) |
| DeepSeek accuracy on TruthfulQA MC1 | ||
| Short context | 33.3% ( ) | — |
| Full context | Condensed | |
| ( 5,000 chars) | ( 5,000 chars) | |
| Sub-call count | 28 (84.8%) | 5 (15.2%) |
| MC choices preserved | 27/28 (96.4%) | 2/5 (40.0%) |
| Mean prompt length (chars) | 15,400 | 1,800 |
| Typical prefix category | extraction | refinement/summary |
| Cases with errors | 0/3 cases | 1/2 cases affected |
| Source | df | (%) | (%) | Magnitude | |
|---|---|---|---|---|---|
| Benchmark | 3 | 3,904 | 15.1 | 15.1 | Large |
| Model benchmark | 14 | 104 | 1.9 | 1.9 | Small |
| Model | 5 | 234 | 1.5 | 1.5 | Small |
| Scaffold benchmark | 9 | 126 | 1.5 | 1.5 | Small |
| Model scaffold | 15 | 25 | 0.5 | 0.5 | Negligible |
| Scaffold | 3 | 119 | 0.5 | 0.5 | Negligible |
| Model | (%) | Avg range (pp) | Max range (pp) | Overall safe (%) |
|---|---|---|---|---|
| DeepSeek V3.2 | 2.1 | 27.7 | 47.5 | 66.8 |
| Mistral Large 2 | 0.2 | 17.3 | 25.6 | 69.7 |
| Opus 4.6 | 2.5 | 16.2 | 24.7 | 80.5 |
| GPT-5.2 | 0.4 | 14.0 | 23.9 | 70.5 |
| Llama 4 Maverick | 0.1 | 13.6 | 19.0 | 68.8 |
| Gemini 3 Pro | 0.1 | 6.1 | 12.8 | 89.2 |
| Hyp. | Comparison | RD (pp) | Wald 95% CI | Wald 90% CI | Bootstrap 95% CI |
|---|---|---|---|---|---|
| H1a | ReAct vs. direct | [ , ] | [ , ] | [ , ] | |
| H1b | Multi-agent vs. direct | [ , ] | [ , ] | [ , ] | |
| H1c | Map-reduce vs. direct | [ , ] | [ , ] | [ , ] |
| Benchmark | ReAct | Multi-agent | Map-reduce |
|---|---|---|---|
| BBQ | [ , ] † | [ , ] | [ , ] |
| Sycophancy | [ , ] † | [ , ] | [ , ] |
| TruthfulQA | [ , ] † | [ , ] | [ , ] |
| XSTest/OR-Bench | [ , ] | [ , ] † | [ , ] |