cs.AIMar 8, 2026

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety

Authors: David Gringras

Organizations: Harvard University

Abstract

Safety benchmarks usually test "bare" models that receive prompts and output responses, but real-world deployments "wrap" those models in complex scaffolds. How much do these scaffolds affect model safety as measured by benchmarks? We test six leading models on four pre-registered safety benchmarks with a direct API and three scaffolds: ReAct, multi-agent, and map-reduce. We conducted 60,112 scored evaluations. On average, how safety is measured matters more than scaffolding does: we find that using a multiple choice vs. open-ended format for otherwise-identical benchmark items changes measured safety by about 5-20 percentage points (pp). The two formats are scored with different methods (answer extraction and an LLM judge), so the gap is due to measurement rather than differences in latent safety. Using a heuristic to classify model refusals would have led to different findings in four of five cases. Benchmark choice explains 15.1% of the variation in outcomes; scaffold architecture explains 0.5%, about 33x less. We find that map-reduce scaffolds, a form of structure-destroying delegation that strips answer options by decomposing prompts, reduce pooled measured safety by 7.3 pp (95% CI: 6.4 to 8.1). The pooled effects for ReAct and multi-agent scaffolds are within our pre-registered +/-2 pp margin of equivalence. However, there are large differences across models for specific benchmarks and scaffolds that are hidden by pooled estimates: for example, on the same sycophancy benchmark items, Opus 4.6 has 16.8 pp lower measured safety with a map-reduce scaffold, while Llama 4 has 18.8 pp higher measured safety. Composite reliability is G = 0.251 (95% CI: [0.000, 0.879]). This wide confidence interval, which spans "of little use" to "very good", does not support using a single composite measure of model safety as the basis for go/no-go decisions about model deployment.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. How Sensitive Are Safety Benchmarks to Judge Configuration Choices?

    Apr 27, 2026Xinran ZhangSafety Benchmarks

  2. Evaluation Awareness Is Not One Capability: Evidence from Open Language Models

    Jun 22, 2026Nilesh Nayan, Aishwarya Sampath Kumar, Rishiraj Girmal +5Safety Benchmarks

  3. Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety

    May 21, 2026Piercosma Bisconti, Matteo Prandi, Federico Pierucci +11Large Language Model SafetySafety Benchmarks