Do Generative Priors Align with Human Naturalness Perception?
Organizations: Communication Science Laboratories, NTT, Inc.
Abstract
Visual generative models are trained to capture the probability distributions of natural images, yet whether their native priors reflect the regularities governing human perception of image naturalness remains an open question. Here, we probe these priors through native prediction errors across 25 open image and video generators. Because raw single-image losses are dominated by scene content and visual complexity, we evaluate directional loss differences using content-preserving, paired relational interventions that selectively disrupt facial configurations or physical illumination consistency while limiting changes in low-level image statistics. Across both domains, these loss differences reproduce human-like selective sensitivities and tolerances, capturing the classic Thatcher effect on faces and shape-dependent responses to illumination inconsistencies. Notably, these loss differences reliably track continuous gradations of human naturalness judgments across individual stimulus pairs (peaking at on faces and on physical scenes) and retain unique human-aligned signals even after controlling for feature distances from frozen vision encoders and standard image quality metrics. We also find that while overall sensitivity to these violations broadly covaries with human alignment across models, the two systematically decouple along denoising schedules, with alignment peaking earlier than sensitivity, revealing that human-like naturalness judgments dissociate from generic violation detection. Together, these findings demonstrate that learning visual distributions yields generative loss landscapes that capture distinct aspects of human naturalness perception.
Figures & tables
Appendix figures & tables39 assets
Supplementary material from the paper’s appendix.
Appendix
| Family | Model | Identifier / Source Repository | Architecture | Evaluated Resolution | Steps ( ) | Noise ( ) |
| Pixel-space image models | ||||||
| JiT | JiT-B/32 | LTH14/JiT::jit-b-32/checkpoint-last.pth | Pixel DiT (ImageNet) | 100 | 20 | |
| JiT | JiT-L/32 | LTH14/JiT::jit-l-32/checkpoint-last.pth | Pixel DiT (ImageNet) | 100 | 20 | |
| JiT | JiT-H/32 | LTH14/JiT::jit-h-32/checkpoint-last.pth | Pixel DiT (ImageNet) | 100 | 20 | |
| PixelGen | PixelGen 512 | zehongma/PixelGen::PixelGen_XXL_T2I.ckpt | Pixel DiT / Flow | 100 | 20 | |
| HiDream | HiDream-O1-Image | HiDream-ai/HiDream-O1-Image | Pixel-space UiT / Flow | 100 | 20 | |
| Encoder | Thatcher | Illumination |
|---|---|---|
| ResNet-50 | (output) | |
| CLIP ViT-B/16 | ||
| DINOv2-Base | (output) | |
| DINOv2-g/14 | ||
| DINOv3-H+/16 | (output) | |
| DINOv3-7B/16 |
| Pool | Predictor | Identifier / checkpoint | Training objective | Params. (B) | Readout / score |
| Encoder | ResNet-50 (ImageNet-1K, torchvision V2) | torchvision::IMAGENET1K_V2 | ImageNet supervision | .024 | Block spatial mean; cosine distance |
| Encoder | CLIP ViT-B/16 T,I | openai/clip-vit-base-patch16 | Image–text contrastive | .086 | Unprojected block CLS; cosine distance |
| Encoder | DINOv2-Base | facebook/dinov2-base | Vision-only self-supervision | .087 | Corresponding patch-token distance |
| Encoder | SigLIP 2 Base/16 I (224) | google/siglip2-base-patch16-224 | Image–text sigmoid contrastive | .093 | LayerNorm, attention pooling; cosine distance |
| Encoder | PE-Core B/16 I (224) | facebook/PE-Core-B16-224 | Vision–language representation | .094 | Block spatial mean; cosine distance |
| Encoder | DINOv2-g/14 spatial | facebook/dinov2-giant | Vision-only self-supervision | 1.136 | Corresponding patch-token distance |
| ID | Evaluated model | B | Published entry | Score |
| Image models: Artificial Analysis human-preference Elo | ||||
| P05 | HiDream-O1-Image | 8.805 | HiDream-O1-Image | 979 |
| L01 | Stable Diffusion v1.5 | 0.86 | Stable Diffusion 1.5 | 468 |
| L02 | SDXL Base 1.0 | 2.6 | Stable Diffusion XL 1.0 | 688 |
| L03 | Stable Diffusion 3 Medium | 2 | Stable Diffusion 3 Medium | 726 |
| L04 | FLUX.1 dev | 12 | FLUX.1 [dev] | 842 |
| Models ( ) | Task | |||
|---|---|---|---|---|
| Image (11) | Thatcher | |||
| Illumination | ||||
| Video (7) | Thatcher | |||
| Illumination |
| Partial Spearman | ||||
|---|---|---|---|---|
| Models | Task | |||
| Image | Thatcher | 11 | ||
| Illumination | 11 | |||
| Video | Thatcher | 7 | ||
| Illumination | 7 | |||
| Condition | Image models | Video models |
|---|---|---|
| Thatcher (upright) | A close-up frontal photograph of a person’s face, shown upright. | A static camera shot showing a close-up frontal view of a person’s face, shown upright. |
| Thatcher (inverted) | An upside-down close-up frontal photograph of a person’s face. | A static camera shot showing an upside-down close-up frontal view of a person’s face. |
| Illumination (all conditions) | A computer-generated 3D scene of multiple geometric objects on a flat tabletop against a dark background. | A static camera shot of a computer-generated 3D scene with multiple geometric objects on a flat tabletop against a dark background. |
| Thatcher | Illumination | |||||
|---|---|---|---|---|---|---|
| Model | Empty | Desc. | [95% CI] | Empty | Desc. | [95% CI] |
| SD v1.5 | ||||||
| SDXL | ||||||
| SD3 | ||||||
| FLUX.1 dev | ||||||
| PixelGen | ||||||
| Task | Same-half | Cross-half | |
|---|---|---|---|
| Thatcher | |||
| Illumination |