cs.AISep 30, 2026

Frozen Scenes, Shifting Winners: Configuration Fragility in Text-to-3D Evaluation

Authors: Anson Y. Lam, Shuqing Li, Michael R. Lyu

Organizations: Department of Computer Science and Engineering The Chinese University of Hong Kong Hong Kong, China

Abstract

Can a text-to-3D leaderboard change when every generated scene stays fixed? We audit this question for rendered-image evaluation, where camera settings and caption wording become part of the measurement protocol. Across 300 frozen scenes from six generators, we vary eight render and caption factors for 19 alignment evaluators plus one perceptual-quality control, then test four targeted scene degradations. Peak configuration variance exceeds between-generator variance for 17/19 alignment evaluators, with prompt-bootstrap lower bounds above 1 for 11/19. Rankings are more stable than scores, yet 18/19 evaluators change their point-estimate winner under some configuration. Pairwise protocol margin envelopes show which comparisons keep their direction across the tested settings. Selected pairs have opposite pointwise intervals, but no reversal survives simultaneous inference over the full search. Thus the observed winner changes are descriptive, not confirmed changes in generator superiority. Sensitivity remains separate: no evaluator, even the prompt-free control, exceeds 67% tie-adjusted directional discrimination on layout scrambling, which is diagnostic rather than human-validated ground truth. The audit separates score stability, decision uncertainty, and targeted sensitivity, and recommends reporting (generator, score, card ID) with protocol-dependent comparisons and selection-aware uncertainty.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation

    Oct 1, 2026Yu Huang, Jungang Li, Zhiyuan Wang +8Text-To-Video Generation ModelVideo Generation

  2. DynEval: Holistic Evaluations of T2I Generative Models in the Wild

    Jul 13, 2026Shyam Marjit, Dheeraj Baiju, Anuj Shikarkhane +3Modern Text-To-Image ModelsText-To-Image

  3. 3D-DefectBench: A Controlled Factorial Study of Vision-Language Model Evaluation Pipelines for Fine-Grained 3D Generation Defects

    Jul 12, 2026Zhenyu Zhao, Nanshan Jia, Jihyeon Je +7Controlled Study