Don't Read the Log: Execution Traces Contaminate Verifiers in Video-Generation Agents
Organizations: RIKEN
Abstract
Agentic video-generation systems close a loop between a generator and a verifier: an LLM plans shots, calls a text-to-video model, and a multimodal judge decides whether the result satisfies the request. To diagnose where a long workflow fails, recent harnesses deliberately show the judge more than the video-the agent's execution trace, its plan, the narration it synthesized. We ask whether this auxiliary text moves the judge's verdict on purely \emph{visual} requirements, holding the frames fixed. On a benchmark of 109 generated two-event clips with manual labels, in which the requested event is either visibly completed or visibly missing, a trace that reports a successful tool call makes three open-weight Qwen-VL judges (7B, 8B, 32B) accept -- of the failures, up from -- without text, and a contradicting trace makes them reject up to of correct clips; an instruction to ``use only the frames'' does not remove the effect. Frontier closed judges are essentially unmoved on the same clips, showing that the vulnerability is a property of the judge's learned trust in tool logs rather than of the task. Plan-derived text carries no clip-specific information, so it can only shift a judge's operating point, and in a repair loop that shift becomes a cap on the true pass rate that no repair policy can exceed; the cap matches simulation to two decimals. In the loop, contamination is exploited without any adversarial agent: an honest LLM planner that always regenerates ends with a judge pass rate of and a human-labelled pass rate of , and a pipeline in which a cheap checker writes its verdict into the trace launders that checker's errors into a stronger final judge ( false accepts).
Figures & tables
| open-weight judges | frontier closed judges | |||||
| auxiliary text | Qwen2.5-VL-7B | Qwen3-VL-8B | Qwen3-VL-32B | GPT-5.4-mini | GPT-5.5 | Claude Opus 5 |
| none | 0.96 / 0.17 | 0.96 / 0.19 | 0.82 / 0.07 | 0.98 / 0.12 | 0.88 / 0.09 | 0.88 / 0.10 |
| transcript, supports | 1.00 / 0.37 | 0.98 / 0.47 | 0.98 / 0.17 | 0.98 / 0.15 | 0.91 / 0.09 | 0.94 / 0.10 |
| subtitle, supports | 0.92 / 0.19 | 0.98 / 0.54 | 0.82 / 0.10 | 1.00 / 0.27 | 0.88 / 0.09 | 0.88 / 0.10 |
| log, supports | 1.00 / 0.78 | 1.00 / 0.90 | 1.00 / 0.83 | 0.98 / 0.14 | 0.88 / 0.09 | 0.88 / 0.10 |
| + visual-only prompt | 1.00 / 0.63 | 0.96 / 0.54 | 1.00 / 0.41 | 1.00 / 0.12 | 0.85 / 0.06 | 0.88 / 0.10 |
| judge | no memory | memory: asset locked | memory: re-cast |
|---|---|---|---|
| Qwen3-VL-8B | 1.00 / 0.47 | 1.00 / 1.00 | 0.88 / 0.29 |
| Qwen3-VL-32B | 1.00 / 0.29 | 1.00 / 0.47 | 0.00 / 0.00 |
| GPT-5.4-mini | 1.00 / 0.18 | 1.00 / 0.12 | 1.00 / 0.12 |
| Claude Opus 5 | 1.00 / 0.00 | 1.00 / 0.00 | 0.76 / 0.00 |
| judge | evidence | planner | judge pass | true pass | hacked | cost/ep | round-0 |
|---|---|---|---|---|---|---|---|
| Qwen3-VL-8B | Naive | always regenerate | 1.00 | 0.28 | 0.72 | 0.36 | 0.72 |
| best-of- | 1.00 | 0.28 | 0.72 | 0.28 | |||
| cost-greedy | 0.75 | 0.22 | 0.72 | 0.30 | |||
| LLM planner | 1.00 | 0.28 | 0.72 | 0.36 | |||
| LP | always regenerate | 1.00 | 0.86 | 0.14 | 1.11 | 0.14 | |
| best-of- | 1.00 | 0.86 | 0.14 | 0.86 |
| judge | evidence | ceiling | best observed true pass | |
|---|---|---|---|---|
| Qwen3-VL-8B | Naive | 0.72 | 0.28 | 0.28 |
| Qwen3-VL-8B | LP | 0.14 | 0.86 | 0.86 |
| Qwen3-VL-32B | Naive | 0.25 | 0.75 | 0.75 |
| Qwen3-VL-32B | Naive + 8B check | 0.69 | 0.31 | 0.31 |
| Qwen3-VL-32B | LP | 0.08 | 0.92 | 0.92 |
| naive | OCR-mask | |||
| subtitle | pres. | abs. | pres. | abs. |
| none | 0.97 | 0.14 | 0.97 | 0.14 |
| supports | 0.97 | 0.50 | 0.97 | 0.14 |
| contradicts | 0.82 | 0.06 | 0.97 | 0.17 |
| evidence-first, supports | 0.77 | 0.28 | ||
| evidence-first, contradicts | 0.50 | 0.06 | ||
| naive | OCR-mask | |||
| subtitle | pres. | abs. | pres. | abs. |
| none | 0.97 | 0.14 | 0.97 | 0.14 |
| supports | 0.97 | 0.50 | 0.97 | 0.14 |
| contradicts | 0.82 | 0.06 | 0.97 | 0.17 |
| evidence-first, supports | 0.77 | 0.28 | ||
| evidence-first, contradicts | 0.50 | 0.06 | ||
| requirement | Naive | LP |
|---|---|---|
| plan written before generation | 0.93 | 0.89 |
| narration track added | 1.00 | 1.00 |
| generator called with 81 frames | 1.00 | 1.00 |
| no tool call returned an error | 1.00 | 1.00 |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| id | (requirement) | filler | pres. | abs. | |
| mug | picks up a red mug | drinks from the mug | holds it, looks out of the window | 2 | 1 |
| door | opens the door | walks through the doorway | stands in the doorway | 2 | 4 |
| hat | puts on a yellow hat | waves at the camera | adjusts the brim | 2 | 4 |
| bench | sits on a bench | opens a newspaper | sits with hands on knees | 2 | 2 |
| apple | picks up a green apple | takes a bite | inspects the apple | 2 | 3 |
| book | picks up a book | opens it and reads | holds the closed book | 2 | 0 |
| Qwen2.5-VL-7B | Qwen3-VL-8B | Qwen3-VL-32B | ||||
|---|---|---|---|---|---|---|
| auxiliary text | true | distr. | true | distr. | true | distr. |
| none | 0.89 | 0.15 | 0.82 | 0.10 | 0.80 | 0.11 |
| transcript, supports | 0.98 | 0.19 | 0.93 | 0.11 | 0.83 | 0.14 |
| log, supports | 1.00 | 0.35 | 0.98 | 0.17 | 0.88 | 0.16 |
| subtitle, supports | 0.99 | 0.43 | 0.98 | 0.20 | 0.92 | 0.17 |
| transcript, contradicts | 0.31 | 0.15 | 0.33 | 0.10 | 0.25 | 0.10 |