A Matched-Budget Audit Framework for Recaptioned Image-Text Supervision Distributions
Organizations: Seoul National University
Abstract
Recaptioned image-text corpora are now standard for text-to-image (T2I) training, with vision--language model (VLM) captioners replacing sparse alt-text by dense descriptions. A recaptioned corpus is a supervision distribution induced by a documented captioning policy (), captioner (), and source corpus (). Length-correlated proxies miss caption-register artifacts and downstream T2I benchmarks entangle the corpus with training choices, so this distribution is hard to audit at corpus scale. We introduce a reusable matched-budget audit framework for recaptioned supervision distributions : at a fixed text budget of it reports a five-axis profile spanning prompt-side coverage, image-conditioned faithfulness, and caption-surface health, with claimed controllable basic units (CBU) as the common claim unit. We instantiate the framework on seven paired comparisons over five public source corpora. Across the four cross-corpus pairs, the released surface raises supported CBU per caption by to under both Qwen and Gemma Judges, and on CC12M the same framework exposes a long-vs-dense frontier that is consistent across both judges and four budgets. We release the audited multi-source recap corpus ( 490M) together with the audit-artifact bundle.
Figures & tables
| Work | Target | Natural | Corpus | Unit | Image | Budget | Prompt |
| REVISE ( 56 ) | visual datasets | ✓ | ✓ | – | – | – | – |
| LAION’s Den ( 4 ) | image–alt-text pairs | ✓ | ✓ | – | – | – | – |
| Hirota et al. ( 23 ) | caption enrichment | ✓ | ✓ | ✓ | ✓ | – | – |
| Brack et al. ( 6 ) | training captions | ✓ | ✓ | – | – | – | – |
| TIFA / DSG ( 25 ; 12 ) | generated images | – | – | ✓ | ✓ | – | – |
| FAITHSCORE ( 28 ) | VLM answers | ✓ | – | ✓ | ✓ | – | – |
| Source family | Original supervision | Ours scale | Paired reference surface(s) |
| Photorealistic / web | |||
| DataComp ( 16 ) | web image–text pairs | M | Recap-DataComp ( 32 ) |
| CC12M ( 10 ) | web alt-text | M | CC12M-LLaVA-NeXT ( 8 ) , PixelProse ( 49 ) , CC12M-Qwen3-VL † ( 53 ) |
| LAION-pop ( 46 ) | web alt-text | M | LAION-pop-Llama ( 9 ) |
| PD12M ( 36 ) | Florence-2 + metadata | M | PD12M released ( 50 ) |
| CommonCatalog ( 19 ) | BLIP-2 captions | M | — |
| Source | Reference release | Avg lex | Opener | Top-100 raw | Top-100 content | Distinct-3 |
| DataComp | Recap-DataComp ( 32 ) | |||||
| CC12M | CC12M-LLaVA-NeXT ( 8 ) | |||||
| PixelProse ( 49 ) | ||||||
| CC12M-Qwen3-VL † ( 53 ) | ||||||
| LAION-pop | LAION-pop-Llama ( 9 ) | |||||
| PD12M | PD12M released ( 50 ) |
| Axis | Property | Input | Metric | Rows |
| Text budget | Coverage | Avg. lex, -eligibility | M | |
| Prompt-pool support | Coverage | , pools | prompt-mass support , -gram JSD | k |
| Claimed density | Coverage | CBU/cap , CBU/100 lex | k | |
| Surface concentration | Health | top-100 prefix mass , distinct-3 | M | |
| Support and risk | Faithfulness | , | k |
| Qwen Judge | Gemma Judge | ||||||
| Dataset | Avg lex | CBU/cap | Pool-wins | Sup. CBU/cap | Risk | Sup. CBU/cap | Risk |
| DataComp | |||||||
| LAION-pop | |||||||
| PD12M | |||||||
| Danbooru | |||||||
| Surface stats | Qwen Judge | Gemma Judge | |||||
| Dataset | Surface | CBU/cap | CBU/100lex | Sup. CBU/cap | Risk | Sup. CBU/cap | Risk |
| CC12M | Ours | ||||||
| Naive Qwen3.5-35B-A3B | |||||||
| DataComp | Ours | ||||||
| Naive Qwen3.5-35B-A3B | |||||||
| Metric | Ours | Naive | Naive Ours | Long-form refs. |
| Mean lex | – | |||
| Lex overflow 248 | pp | – | ||
| Top-100 raw prefix mass | – | |||
| Top-100 content prefix mass | – | |||
| Distinct-3-gram rate | – |
| Surface stats | Qwen Judge | Gemma Judge | ||||
| Surface | CBU/cap | CBU/100lex | Sup. CBU/cap | Risk | Sup. CBU/cap | Risk |
| Ours | ||||||
| CC12M-LLaVA-NeXT ( 8 ) | ||||||
| PixelProse ( 49 ) | ||||||
| CC12M-Qwen3-VL † ( 53 ) | ||||||
| (a) Judge–human agreement | |||||
| Judge | Overall | Ours | Pooled refs. | ||
| Qwen | |||||
| Gemma | |||||
| (b) Human image support | |||||
| Surface | Yes | Uncertain | Explicit no | Other | |
| Ours | |||||
| (a) DataComp text-space probes | ||||||
| Vendi | eRank | Coverage | ||||
| Encoder | Ours | Ref | Ours | Ref | Ours | Ref |
| Qwen3-Emb-4B | ||||||
| Qwen3-Emb-8B | – | – | ||||
| BGE-M3 | ||||||
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol or term | Meaning |
| Supervision distribution | |
| , | source image corpus and one of its images |
| , | captioner and captioning policy |
| caption–image pairs written by under on ; the audit target | |
| , | caption marginal and caption–image joint |
| (Ours) | the released corpus, the first audit target |
| Audit stage | Input | Requests | H200 GPU-hours |
| ClaimedCBU@64 | text | k | |
| GroundedCBU@64 | text + image | k | tens |
| CBU-VQA | question + image | k | tens |
| Pool name | Target | Source / hosting |
| civitai_flux_prompts_aconexx | FLUX | https://huggingface.co/datasets/Aconexx/CivitAI-Flux-Prompts |
| flux_improved_k_mktr | FLUX | https://huggingface.co/datasets/k-mktr/improved-flux-prompts |
| flux_prompts_chrisgoringe | FLUX | https://huggingface.co/datasets/ChrisGoringe/flux_prompts |
| flux_prompts_regpeter | FLUX | https://huggingface.co/datasets/regpeter/flux_prompts |
| sd_prompts_2m_andyyang | Stable Diffusion | https://huggingface.co/datasets/andyyang/stable_diffusion_prompts_2m |
| sdxl_refiner_prompts_falah | SDXL refiner | https://huggingface.co/datasets/Falah/1M_SDXL_Refiner_Prompts |
| Dataset | Surface | -elig | Responses | Questions | Q/cap | Risk |
| DataComp | Ours | |||||
| DataComp | Recap-DataComp ( 32 ) | |||||
| Danbooru | Ours | |||||
| Danbooru | Danbooru-Florence ( 60 ; 29 ) | |||||
| LAION-pop | Ours | |||||
| LAION-pop | LAION-pop-Llama ( 9 ) |
| Risk | ||||||||
| Surface | Aligned | Valid | Cap | Zero | VQA | Q | Qwen | Gemma |
| Ours | ||||||||
| CC12M-LLaVA-NeXT | ||||||||
| PixelProse | ||||||||
| CC12M-Qwen3-VL † | ||||||||
| Source | Pair | Surface | Opener | Top-100 raw | Top-100 content | Distinct-3 |
| DataComp | Recap-DataComp | Ref | ||||
| Ours | ||||||
| CC12M | CC12M-LLaVA-NeXT | Ref | ||||
| Ours | ||||||
| PixelProse | Ref | |||||
| Ours |
| Sup. CBU/cap | |||||
| Dataset | Judge | all types | six types | Risk, six types | |
| DataComp | Qwen | ||||
| Gemma | |||||
| LAION-pop | Qwen | ||||
| Gemma | |||||
| PD12M | Qwen | ||||
| Qwen Judge | Gemma Judge | ||||||||
| Ours | Refs | Ours | Refs | ||||||
| Type | Sup. | Risk | Sup. | Risk | Sup. | Risk | Sup. | Risk | Agree. |
| object | |||||||||
| attribute | |||||||||
| relation | |||||||||
| count | |||||||||
| Qwen Judge | Gemma Judge | ||||
| Dataset | Surface | Sup CBU/cap | Risk | Sup CBU/cap | Risk |
| CC12M | Ours | ||||
| CC12M-LLaVA-NeXT | |||||
| PixelProse | |||||
| CC12M-Qwen3-VL † | |||||
| DataComp | Reference | ||||
| Caption screen | Image support | |||||||
| Surface | Licensed | Atomic | Type | Yes | Uncertain | No | Other | |
| Ours | ||||||||
| LLaVA-NeXT | ||||||||
| Qwen3-VL-8B | ||||||||
| PixelProse | ||||||||
| Claim type | Ours | LLaVA-NeXT | Qwen3-VL-8B | PixelProse |
| Attribute | – | |||
| Camera | ||||
| Count | ||||
| Lighting | – | |||
| Object | ||||
| Relation |
| Qwen Judge | Gemma Judge | ||||||
| Dataset | Surface | VQA | Q | Sup CBU/cap | Risk | Sup CBU/cap | Risk |
| CC12M | Ours | ||||||
| Naive | |||||||
| Naive (greedy) | |||||||
| DataComp | Ours | ||||||
| Naive | |||||||
| Vendi | eRank | Coverage@10 | Density@10 | kNN cos | ||||||
| Encoder | Ours | Ref | Ours | Ref | Ours | Ref | Ours | Ref | ||
| Qwen3-Emb-4B | ||||||||||
| Qwen3-Emb-8B | – | – | – | – | ||||||
| BGE-M3 | ||||||||||
| LongCLIP tokens / cap | I2T | T2I | |||||
| Surface | mean | p95 | trunc. | R@1 | R@5 | R@1 | R@5 |
| Ours | |||||||
| CC12M-LLaVA-NeXT ( 8 ) | |||||||
| PixelProse ( 49 ) | |||||||
| CC12M-Qwen3-VL † ( 53 ) | |||||||
| Naive Qwen3.5-35B-A3B | |||||||
| LongCLIP tokens / cap | I2T | T2I | |||||
| Surface | mean | p95 | trunc. | R@1 | R@5 | R@1 | R@5 |
| Ours | |||||||
| CC12M-LLaVA-NeXT | |||||||
| PixelProse | |||||||
| CC12M-Qwen3-VL † | |||||||
| Naive Qwen3.5-35B-A3B | |||||||
| Source | Surface | Mean tokens | CLIP-77 trunc | LongCLIP-248 trunc | SigLIP2-64 trunc |
| DataComp | Ours | ||||
| Recap-DataComp | |||||
| CC12M | Ours | ||||
| Naive Qwen3.5-35B-A3B | |||||
| CC12M-LLaVA-NeXT | |||||
| PixelProse | — | — |
| Group | Fields | Role |
| Source keys | source_dataset , source_url , source_url_norm , source_url_sha1 , source_sha256 | source join |
| Image locator | asset_instance_id , image_shard , image_member | image binding |
| Caption | caption_text , caption_sha256 , caption_record_id | released text |
| Generation | caption_model_family , caption_model_artifact , prompt_tokens , completion_tokens | provenance |
| Result | Scope | Artifact |
| Claim yield (Table 5 ) | cross-corpus pairs | all_cbu_b64_summary.csv |
| Support and risk (Tables 5 , 8 , Fig. 2 ) | both Judges | cbu_vqa_by_category_b64.json |
| Length and prefix statistics (Tables 3 , 5 ) | caption text | cpu_text_metrics/ |
| Pool comparisons (Figs. 1 , 3 ) | seven prompt pools | prompt_support_bootstrap_b64_n2_250k_2026-04-24.tsv |
| Claim yield and budget sweep (Table 8 , Fig. 2 ) | CC12M | cc12m_budget_frontier_plot.csv |
| Policy control (Tables 6 , 7 ) | CC12M, DataComp | naive_qwen35_*/ |
| Stage | Count | Share |
| Target URLs | — | |
| Processed | ||
| Download-stage survival | ||
| Usable images | ||
| Top failure families: | ||
| image_too_small | ||