cs.CLMar 12, 2026

Cross-Context Review: Improving LLM Output Quality by Separating Production and Review Sessions

Authors: Tae-Eun Song

Organizations: Daejeon Jungang Cheonggua Co., Ltd.

Abstract

Large language models struggle to catch errors in their own outputs when the review happens in the same session that produced them. This paper introduces Cross-Context Review (CCR), a straightforward method where the review is conducted in a fresh session with no access to the production conversation history. We ran a controlled experiment: 30 artifacts (code, technical documents, presentation scripts) with 150 injected errors, tested under four review conditions -- same-session Self-Review (SR), repeated Self-Review (SR2), context-aware Subagent Review (SA), and Cross-Context Review (CCR). The central result is that a second review helps only when it happens in a fresh session: CCR (F1 28.6%) outperforms a second review in the same session (SR2, 21.7%) robustly, both in the first run (paired t, p<0.001) and in the three-run average (Holm-adjusted p=0.004). This version updates the broader comparisons. Averaged across runs, and excluding one SR run whose records could not be verified, CCR is not significantly ahead of context-aware subagent review (SA, 23.8%; p=0.057) or of a single same-session review (SR, 27.1%; p=0.26); the first version's advantages over these two baselines came from run 1. CCR needs no infrastructure and costs one extra session.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. When Does a Second Model Help? Cross-Model Review in LLM Verification

    Oct 1, 2026Tae-Eun SongPeer ReviewProvenance

  2. Cross-Model LLM Code Review: Should you use Claude to review Codex or vice versa?

    Jul 22, 2026Zuodong Xiang, Yike Zhang, YueMing Zhang +1Code QualityLlm-As-A-Judge

  3. CRJudgeBench: Can AI Detect Plausible but Invalid Code Reviews?

    Sep 29, 2026Yue Pan, Jiawei Li, Ziyuan Zhang +2Code QualityRaw Judge Outputs