cs.CVOct 1, 2026

A Matched-Budget Audit Framework for Recaptioned Image-Text Supervision Distributions

Authors: Giyeong Oh, Junghun Park, Yuhan Bae, Youngjae Yu

Organizations: Seoul National University

Abstract

Recaptioned image-text corpora are now standard for text-to-image (T2I) training, with vision--language model (VLM) captioners replacing sparse alt-text by dense descriptions. A recaptioned corpus is a supervision distribution induced by a documented captioning policy (ππ), captioner (VcV_c), and source corpus (CC). Length-correlated proxies miss caption-register artifacts and downstream T2I benchmarks entangle the corpus with training choices, so this distribution is hard to audit at corpus scale. We introduce a reusable matched-budget audit framework for recaptioned supervision distributions Dπ,Vc,CD_{π,V_c,C}: at a fixed text budget of B=64B = 64 it reports a five-axis profile spanning prompt-side coverage, image-conditioned faithfulness, and caption-surface health, with claimed controllable basic units (CBU) as the common claim unit. We instantiate the framework on seven paired comparisons over five public source corpora. Across the four cross-corpus pairs, the released surface raises supported CBU per caption by +3.39+3.39 to +6.36+6.36 under both Qwen and Gemma Judges, and on CC12M the same framework exposes a long-vs-dense frontier that is consistent across both judges and four budgets. We release the audited multi-source recap corpus (≈\approx 490M) together with the audit-artifact bundle.

Figures & tables

Appendix figures & tables21 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. CAPEval: A Decoupled Caption Evaluation across Understanding and Generation

    Aug 3, 2026Zhipeng Liu, Haochen Wang, Zhaoxiang ZhangCaptionsMultimodal Understanding

  2. A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions

    Jul 25, 2026Zhijiang Tang, Jiaxin Qi, Kaihua Tang +2Image Difference Captioning