Atomic Visual Entailment: Enhancing Zero-Shot Vision-Language Reasoning through Atomic Fact Decomposition and Learned Selection
Organizations: Leiden Institute of Advanced Computer Science (LIACS), Leiden University, The Netherlands
Abstract
Visual entailment (VE) asks whether an image supports, contradicts, or leaves undecided a textual hypothesis. Strong results come from fine-tuning large vision-language models on labelled data, while zero-shot and hybrid approaches remain far behind. A VE hypothesis often bundles several visual claims, yet existing zero-shot methods reason over it as a single unit. We propose Atomic Visual Entailment (AVE), which decomposes the hypothesis into atomic facts, produces candidate predictions from both the full hypothesis and its facts using frozen vision-language models, and predicts the final label with a lightweight classifier trained only on how those candidates behave. We find that decomposition helps only when the hypothesis context is preserved: judging facts in isolation is worse than not decomposing at all. Full-hypothesis and atomic prediction make complementary errors, and learning which to trust recovers far more of that complementarity than majority voting, reaching 0.803 test accuracy on SNLI-VE without fine-tuning any vision-language model. AVE also localises the visual evidence behind its prediction without region-level supervision. These results suggest that learning which candidate prediction to trust can close much of the gap to fine-tuned systems, offering a practical alternative where fine-tuning a vision-language model directly would need more labelled data or compute than is available.
Figures & tables
| Atoms | Share | E | N | C |
|---|---|---|---|---|
| 1 | 63.6 | 39.2 | 27.4 | 33.4 |
| 2 | 31.8 | 23.6 | 42.6 | 33.9 |
| 3 | 3.9 | 20.4 | 51.1 | 28.5 |
| 4+ | 0.7 | 25.0 | 49.2 | 25.8 |
| Development | Test | |||
| Method | Acc. | F1 | Acc. | F1 |
| Best full-hypothesis | 0.749 | 0.732 | 0.754 | 0.737 |
| Best atomic | 0.750 | 0.739 | 0.757 | 0.746 |
| Best self-decomposed | 0.762 | 0.757 | 0.765 | 0.760 |
| AVE-MV (Qwen3) | 0.752 | 0.744 | 0.756 | 0.748 |
| AVE-MV (InternVL3) | 0.761 | 0.751 | 0.768 | 0.757 |
| Method | Dev | Test | |
|---|---|---|---|
| Sup. | EVE-Image ( Xie et al., 2019 ) | 0.716 | 0.711 |
| UNITER ( Chen et al., 2020 ) | – | 0.794 | |
| OFA ( Wang et al., 2022 ) | 0.910 | 0.912 | |
| ZS | CLIP ( Song et al., 2022 ) | 0.676 | 0.682 |
| IdealGPT ( You et al., 2023 ) | 0.553 | – | |
| AVE-MV (ours) | 0.765 | 0.769 |
| Entail. | Contra. | Overall | |
|---|---|---|---|
| Matched | 2,173 | 1,933 | 4,106 |
| Coverage | 0.869 | 0.773 | 0.821 |
| Recall@1 | 0.823 | 0.852 | 0.837 |
| Recall@3 | 0.896 | 0.904 | 0.900 |
| Mean IoU@3 | 0.843 | 0.847 | 0.845 |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Decomposer | Faith. | Compl. | Both |
|---|---|---|---|
| Meta-Llama-3.1-8B | 95 | 88 | 88 |
| Qwen3-8B | 96 | 90 | 90 |
| Qwen2.5-32B | 98 | 96 | 96 |
| Feature variant | Features | |
|---|---|---|
| Full-hypothesis | 4 | 129 |
| Simple atomic | 4 | 129 |
| Structured atomic | 4 | 129 |
| Full 12 methods | 12 | 409 |
| HistGB | XGBoost | |||
|---|---|---|---|---|
| Feature variant | Acc. | F1 | Acc. | F1 |
| Full-hypothesis | 0.790 | 0.787 | 0.788 | 0.785 |
| Simple atomic | 0.792 | 0.788 | 0.794 | 0.792 |
| Structured atomic | 0.793 | 0.790 | 0.793 | 0.791 |
| Full 12 methods | 0.798 | 0.796 | 0.801 | 0.798 |
| Parameter | Value |
|---|---|
| objective | multi:softprob |
| eval_metric | mlogloss |
| n_estimators | 1,100 |
| max_depth | 3 |
| learning_rate | 0.025 |
| subsample | 0.88 |