Frozen Scenes, Shifting Winners: Configuration Fragility in Text-to-3D Evaluation
Organizations: Department of Computer Science and Engineering The Chinese University of Hong Kong Hong Kong, China
Abstract
Can a text-to-3D leaderboard change when every generated scene stays fixed? We audit this question for rendered-image evaluation, where camera settings and caption wording become part of the measurement protocol. Across 300 frozen scenes from six generators, we vary eight render and caption factors for 19 alignment evaluators plus one perceptual-quality control, then test four targeted scene degradations. Peak configuration variance exceeds between-generator variance for 17/19 alignment evaluators, with prompt-bootstrap lower bounds above 1 for 11/19. Rankings are more stable than scores, yet 18/19 evaluators change their point-estimate winner under some configuration. Pairwise protocol margin envelopes show which comparisons keep their direction across the tested settings. Selected pairs have opposite pointwise intervals, but no reversal survives simultaneous inference over the full search. Thus the observed winner changes are descriptive, not confirmed changes in generator superiority. Sensitivity remains separate: no evaluator, even the prompt-free control, exceeds 67% tie-adjusted directional discrimination on layout scrambling, which is diagnostic rather than human-validated ground truth. The audit separates score stability, decision uncertainty, and targeted sensitivity, and recommends reporting (generator, score, card ID) with protocol-dependent comparisons and selection-aware uncertainty.
Figures & tables
| Audit question | Result | Scope and interpretation |
|---|---|---|
| Does configuration variance exceed generator separation? | 17/19 | Peak for the six-generator field. |
| Does this persist under prompt resampling? | 11/19 | Bootstrap lower bound for peak exceeds 1. |
| Does the point-estimate winner change? | 18/19 | Across tested factor levels, without claiming a confirmed reversal. |
| Is an opposite pairwise ordering confirmed after search? | 0/19 | Evaluators with a selected reversal under a band over 110,295 contrasts. |
| Which pairwise orders survive all tested protocols? | 4–11/15 | Strict orders per alignment evaluator over a finite set of 387 protocols. |
| Does scrambling lower the score? | Tie-adjusted ceiling across evaluators, including the render-only control. |
| Rank stability | Content acc. (%) | Best view | ||||
|---|---|---|---|---|---|---|
| Evaluator | template | peak | min | #top-1 | ||
| CLIP ViT-B/32 | 2.05 | 3.04 † | 0.75 | 2 | 67 | 1.17 |
| CLIP ViT-B/16 | 2.30 | 2.30 † | 0.64 | 3 | 68 | 1.14 |
| CLIP ViT-L/14 | 3.04 | 3.04 † | 0.70 | 2 | 71 | 1.14 |
| CLIP ViT-L/14@336 | 1.85 | 1.85 † | 0.77 | 2 | 72 | 1.12 |
| OpenCLIP ViT-L/14 | 1.08 | 1.08 | 0.75 | 2 | 70 | 1.10 |
| Slot | Disp. | Choices |
|---|---|---|
| Viewpoint | top-down , bird-eye , (none) | |
| Head noun | rendering , view , image , picture , photo | |
| Scene frame | a virtual scene: , a scene: , scene: , (none) | |
| Connective | depicting , about , of | |
| Article | a , (none) , an |
| Evaluator / factor | Method pair | Levels | Gap [pointwise 95%] | Simultaneous 95% |
|---|---|---|---|---|
| CLIP ViT-B/32 Lighting | Scenethesis I-Design | city night | ||
| CLIP ViT-L/14@336 Pitch | Scenethesis Reason-3D | 15 ∘ 90 ∘ | ||
| MetaCLIP ViT-L/14 Pitch | Scenethesis LayoutVLM | 0 ∘ 75 ∘ | ||
| BLIP-VQA Template | Scenethesis Holodeck | T1 T2 |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Work | Artifact held fixed | Render factors isolated | Caption wording isolated | Decision stability | Controlled degradation | Reporting artifact |
|---|---|---|---|---|---|---|
| AutoMetrics ( Ryan et al., 2026 ) | Equivalent-quality outputs, not one frozen artifact | Not applicable to 3D rendering | Output rephrasing, not a 3D caption wrapper | Metric validity, not generator-winner stability | LLM-generated worse-quality outputs | Metric Cards, not a 3D configuration card |
| T2IScoreScore ( Saxon et al., 2024 ) | Related 2D image families under a fixed prompt | Not applicable to 3D rendering | Semantic error graphs, not caption wrappers | Metric ordering, not generator-winner stability | Semantic error graphs | No analogous card in the stated contribution |
| 3D-DefectBench ( Zhao et al., 2026 ) | Fixed assets across 84 inference designs | Camera protocol and visual input crossed | Prompt schema crossed | Factor effects and judge comparisons for defect detection | Nine human-labeled defect categories | Released data and metadata, not a reporting card |
| Cross-Model VLM-Judge ( Asaria et al., 2026 ) | Fixed mesh pairs across 24 rendered views | Fixed multi-view rig, not render-factor isolation | Fixed judge prompt, not a caption-wording audit | Pairwise asset preference, not generator-winner stability | Face-drop degradation, not room-layout scrambling | Reproducible judge protocol, not a configuration card |
| SceneCritic ( Sengupta et al., 2026 ) | Yes across views and repeated VLM calls | Viewpoint and repeated calls, not 8 OFAT factors | Prompt sensitivity motivated, not experimentally swept | View-dependent method-ranking reversals for one VLM judge | No fixed-artifact degradation audit | SceneOnto and a symbolic critic, not a configuration card |
| T3Bench and Eval3D ( He et al., 2023 ; Duggal et al., 2025 ) | Generated assets, not a same-artifact audit | Many views, not a render-factor audit | Caption or QA pipelines, not a wording audit | Leaderboards, not stability under configurations | Janus or inconsistency probes, not room-layout scrambling | No analogous card in the stated contribution |
| Card field | Required report content |
|---|---|
| Schema and status | card_type , schema_version , evidence status, and the validator-produced card identifier. |
| Scene artifact source | Dataset, generator set, scene and prompt counts, frozen-scene rule, and artifact provenance. |
| Render protocol | Renderer, camera pose, focal length, resolution, lighting, background, random seeds, and post-processing. |
| Caption protocol | Prompt template, paraphrase rule, prompt source, and whether wording was frozen before scoring. |
| View policy | Fixed, averaged, or selected view, declared views, and selection rule. |
| Evaluator identity | Model or metric name, checkpoint/API snapshot, instruction, decoding settings, and call date. |
| Slot | Status | Choices |
|---|---|---|
| Article | optional | a, an |
| Viewpoint | optional | top-down, bird-eye |
| Head noun | optional | image, photo, picture, rendering, view |
| Connective | with a noun | of, about, depicting |
| Scene frame | optional | scene:, a scene:, a virtual scene: |
| Prompt | required | {prompt} |
| Environment map | EV | CCT (K) | Mean -shift |
|---|---|---|---|
| studio | -6.41 | 7802 | -0.067 |
| night | -5.13 | 5051 | -0.173 |
| courtyard | -3.70 | 5759 | -0.017 |
| sunrise | -3.24 | 6004 | -0.044 |
| forest | -2.22 | 7513 | +0.075 |
| interior | -2.20 | 5420 | +0.077 |
| Evaluator | Protocols | Strict ordering | Sign crossing | Tie boundary | Stable winner |
|---|---|---|---|---|---|
| CLIP ViT-B/32 | 387 | 5/15 | 10/15 | 0/15 | none |
| CLIP ViT-B/16 | 387 | 4/15 | 11/15 | 0/15 | none |
| CLIP ViT-L/14 | 387 | 4/15 | 11/15 | 0/15 | none |
| CLIP ViT-L/14@336 | 387 | 5/15 | 10/15 | 0/15 | none |
| OpenCLIP ViT-L/14 | 387 | 5/15 | 10/15 | 0/15 | none |
| OpenCLIP ViT-H/14 | 387 | 5/15 | 10/15 | 0/15 | none |
| Outcome | Count | Percent [95% interval] |
|---|---|---|
| Strict drop | 124 | 41.3 [35.0, 47.7] |
| Exact tie | 77 | 25.7 [20.7, 31.0] |
| Rise | 99 | 33.0 [27.0, 38.7] |
| Tie-adjusted discrimination | 54.2 [48.5, 59.5] |
| Factor | # | Levels |
|---|---|---|
| Resolution (pixel) | 9 | 196, 224, 256, 336, 384, 448, 512 ⋆ , 768, 1024 |
| Focal length (mm) | 7 | 16, 24, 35, 50 ⋆ , 85, 100, 200 |
| Lighting (environment map) | 8 | city ⋆ , courtyard, forest, interior, night, studio, sunrise, sunset |
| Background (RGB) | 10 | 0, 65, 118, 128, 186, 204, 255 ⋆ gray, plus red, green, blue |
| Pitch (degrees) | 7 | 0 ⋆ , 15, 30, 45, 60, 75, 90 |
| Yaw (degrees, at pitch ) | 8 | 0 ⋆ , 45, 90, 135, 180, 225, 270, 315 |
| Group | Members |
|---|---|
| CLIP (OpenAI) | ViT-B/32, ViT-B/16, ViT-L/14, ViT-L/14@336 ( Radford et al., 2021 ) |
| CLIP (open data) | OpenCLIP ViT-L/14, ViT-H/14, ViT-bigG/14 ( Cherti et al., 2023 ) , MetaCLIP ( Xu et al., 2024 ) , DFN ( Fang et al., 2024 ) , EVA-CLIP ( Sun et al., 2023 ) , and CLIPA ( Li et al., 2023b ) |
| CLIP (long-text) | Long-CLIP ( Zhang et al., 2024 ) |
| Sigmoid objective | SigLIP ( Zhai et al., 2023 ) and SigLIP 2 ( Tschannen et al., 2025 ) |
| BLIP-2 ( Li et al., 2023a ) | ITM and ITC heads |
| VQA-based | BLIP-VQA ( Huang et al., 2023 ) and VQAScore ( Lin et al., 2024 ) |
| Evaluator | Res. | Focal | Light | Bg. | Pitch | Yaw | Tmpl. | Para. |
|---|---|---|---|---|---|---|---|---|
| CLIP ViT-B/32 | 7 | 13 | 13 | 8 | 20 | 16 | 28 | 13 |
| CLIP ViT-B/16 | 9 | 12 | 12 | 7 | 18 | 14 | 31 | 12 |
| CLIP ViT-L/14 | 15 | 21 | 20 | 12 | 26 | 23 | 51 | 18 |
| CLIP ViT-L/14@336 | 14 | 16 | 17 | 10 | 24 | 19 | 51 | 18 |
| OpenCLIP ViT-L/14 | 20 | 19 | 21 | 12 | 31 | 26 | 56 | 16 |
| OpenCLIP ViT-H/14 | 15 | 16 | 20 | 14 | 31 | 25 | 60 | 18 |
| Evaluator | Res. | Focal | Light | Bg. | Pitch | Yaw | Tmpl. | Para. |
|---|---|---|---|---|---|---|---|---|
| CLIP ViT-B/32 | 0.34 | 1.34 | 1.09 | 0.46 | 3.04 | 2.01 | 2.05 | 0.97 |
| CLIP ViT-B/16 | 0.36 | 0.88 | 0.86 | 0.24 | 2.07 | 1.35 | 2.30 | 0.72 |
| CLIP ViT-L/14 | 0.57 | 1.35 | 1.05 | 0.35 | 2.10 | 1.37 | 3.04 | 0.78 |
| CLIP ViT-L/14@336 | 0.30 | 0.54 | 0.49 | 0.17 | 1.13 | 0.64 | 1.85 | 0.47 |
| OpenCLIP ViT-L/14 | 0.28 | 0.36 | 0.34 | 0.11 | 0.90 | 0.50 | 1.08 | 0.20 |
| OpenCLIP ViT-H/14 | 0.25 | 0.34 | 0.44 | 0.23 | 1.16 | 0.56 | 2.04 | 0.34 |
| Evaluator | Peak | Prompt 95% CI | Leave-one-gen range | Top-1/2 gap [95% CI] |
|---|---|---|---|---|
| CLIP ViT-B/32 | 3.04 | [1.92, 4.41] | [2.51, 9.06] | +0.91 [+0.12, +1.73] ∗ |
| CLIP ViT-B/16 | 2.30 | [1.51, 3.25] | [1.91, 7.92] | +0.33 [-0.64, +1.28] |
| CLIP ViT-L/14 | 3.04 | [1.97, 4.61] | [2.54, 9.60] | +1.12 [-0.24, +2.38] |
| CLIP ViT-L/14@336 | 1.85 | [1.31, 2.69] | [1.52, 5.94] | +1.71 [+0.46, +2.87] ∗ |
| OpenCLIP ViT-L/14 | 1.08 | [0.76, 1.49] | [0.89, 6.45] | +2.59 [+1.11, +4.03] ∗ |
| OpenCLIP ViT-H/14 | 2.04 | [1.44, 2.84] | [1.70, 12.69] | +1.49 [+0.01, +2.90] ∗ |
| Evaluator | Res. | Focal | Light | Bg. | Pitch | Yaw | Tmpl. |
|---|---|---|---|---|---|---|---|
| CLIP ViT-B/32 | 0.91/1 | 0.89/2 | 0.83/2 | 0.95/1 | 0.75/2 | 0.87/2 | 0.94/1 |
| CLIP ViT-B/16 | 0.97/1 | 0.89/1 | 0.75/3 | 0.97/2 | 0.64/2 | 0.73/1 | 0.89/2 |
| CLIP ViT-L/14 | 0.91/1 | 0.89/1 | 0.70/2 | 0.95/1 | 0.82/2 | 0.94/2 | 0.91/2 |
| CLIP ViT-L/14@336 | 0.87/1 | 0.95/1 | 0.84/2 | 0.97/1 | 0.85/2 | 0.77/1 | 0.89/1 |
| OpenCLIP ViT-L/14 | 0.85/1 | 0.94/1 | 0.75/2 | 0.93/1 | 0.80/2 | 0.89/1 | 0.97/1 |
| OpenCLIP ViT-H/14 | 0.87/1 | 0.95/1 | 0.72/2 | 0.95/1 | 0.81/3 | 0.95/1 | 0.94/1 |
| Evaluator | Keep-half | Biggest-only | Scrambled | Worst-objects | Mean (4) |
|---|---|---|---|---|---|
| CLIP ViT-B/32 | 64 | 74 | 56 | 72 | 67 |
| CLIP ViT-B/16 | 64 | 75 | 57 | 77 | 68 |
| CLIP ViT-L/14 | 69 | 78 | 57 | 80 | 71 |
| CLIP ViT-L/14@336 | 71 | 79 | 58 | 81 | 72 |
| OpenCLIP ViT-L/14 | 66 | 85 | 54 | 76 | 70 |
| OpenCLIP ViT-H/14 | 70 | 86 | 62 | 81 | 75 |
| Evaluator | Resolution | Focal | Lighting | Background | Pitch | Yaw |
|---|---|---|---|---|---|---|
| CLIP ViT-B/32 | 224/1024 | 85/16 | interior/studio | 204/green | 45/0 | 180/45 |
| CLIP ViT-B/16 | 336/224 | 200/16 | forest/night | 255/65 | 45/0 | 180/135 |
| CLIP ViT-L/14 | 448/224 | 200/16 | city/night | 255/green | 30/90 | 90/180 |
| CLIP ViT-L/14@336 | 1024/384 | 200/16 | forest/night | 255/green | 15/90 | 45/180 |
| OpenCLIP ViT-L/14 | 1024/196 | 200/16 | city/night | 128/green | 15/90 | 0/225 |
| OpenCLIP ViT-H/14 | 768/196 | 200/16 | city/night | 128/red | 15/90 | 0/225 |