PRISM: A Category-Theoretic Framework for Measuring and Refining Multimodal Analogies
Organizations: Delft University of Technology
Abstract
Analogical reasoning involves identifying and preserving relational structures across domains. However, existing approaches to AI-driven multimodal analogy generation lack an interpretable measure of whether this structure is understood and maintained in the generated output. We address this gap with Pullback Refinement via Interpretable Structural Mapping (PRISM), a modality-agnostic framework for measuring and improving relational alignment in multimodal analogies, evaluated on visual metaphor generation. PRISM represents analogies as explicit relational mappings grounded in category theory and uses VLMs to instantiate these structures across modalities. Its first component, the pullback score, quantifies relational alignment from the resulting graph representation. On the AnaloBench benchmark, selecting the correct analogy purely by pullback score achieves 82.5% accuracy, demonstrating that the score captures meaningful relational information. PRISM's second component is an iterative refinement loop that uses the pullback score as an in-context feedback signal to iteratively revise the generated image towards greater relational depth. VLM-as-a-judge and human evaluations show that PRISM consistently improves metaphor consistency and analogy appropriateness over zero- shot generation, with human participants preferring the refined output in 57.65% of pairwise comparisons. However, a qualitative analysis reveals that refinement can favour visually crowded compositions rather than genuinely deeper relational correspondences.
Figures & tables
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Correctly generated graphs |
| Qwen 3.5 | 40% |
| Qwen 3.5 with reasoning | 76% |
| Gemma 4 | 78% |
| Gemma 4 with reasoning | 86% |
| Claude Haiku 4.5 | 56% |
| Claude Haiku 4.5 with extended thinking | 84% |
| Metaphor | Visual Elaboration | Generated Image (GPT Image 1.5) |
| Love is a double edged sword. | A gleaming ornate sword suspended in mid-air, its blade splitting into two distinct halves—one radiating warm golden light and blooming roses, the other crackling with cold blue electricity and thorns. The background fades between serene twilight and stormy darkness. Fine mist swirls around the weapon, capturing the duality of beauty and pain. | |
| A sweet tooth is a predator chasing down your smile. | A sleek, shadowy predator with candy-colored fur stalks through a dreamlike landscape toward a luminous, golden smile floating in the distance. The creature’s eyes glow with hunger as it prowls closer, leaving a trail of melting sweets and broken teeth. Soft, surreal lighting contrasts the predator’s dark silhouette against pastel clouds and candy-striped terrain. | |
| The planet is a sinking ship. | A massive Earth sphere tilts precariously in turbulent waters, its continents cracking and flooding. Desperate figures cling to the edges as waves crash over the surface. Smoke rises from fissures. The sky darkens ominously. Lifeboats drift empty nearby, unreachable. The horizon swallows everything in murky depths, evoking apocalyptic urgency and collective doom. | |
| Social media is a hamster wheel. | An exhausted figure endlessly running inside a massive transparent hamster wheel, surrounded by glowing smartphone screens and notification badges. The wheel spins relentlessly in a dimly lit room, casting repetitive shadows. Despite constant motion, the scenery never changes. Scattered digital clutter accumulates around the base as the person runs faster, trapped in an endless cycle of movement without progress. |
| Model | Both Domains Match | Either Domain Matches |
| GPT-Image Zero-Shot | 9.4% | 59.2% |
| GPT-Image Chain-of-Thought | 6.5% | 55.7% |
| Flux-Klein Zero-Shot | 5.9% | 54.2% |
| Stable-Core Chain-of-Thought | 4.8% | 44.0% |
| Stable-Core Zero-Shot | 4.4% | 52.0% |
| Flux-Klein Chain-of-Thought | 4.2% | 52.9% |
| Image Text | Text Image | |||||
| Model | Top-1 (%) | Top-5 (%) | Mean Rank | Top-1 (%) | Top-5 (%) | Mean Rank |
| GPT-Image Zero | 84 | 96 | 1.57 | 82 | 96 | 2.54 |
| GPT-Image CoT | 76 | 91 | 3.19 | 72 | 90 | 3.57 |
| Flux-Klein Zero | 76 | 92 | 2.60 | 67 | 87 | 4.50 |
| Stable-Core Zero | 69 | 87 | 4.22 | 61 | 82 | 6.35 |
| Flux-Klein CoT | 65 | 86 | 4.27 | 59 | 81 | 5.33 |
| Analysis | Condition | Pearson | Spearman | |||
| Concreteness | Flux-Klein Zero | 231 | ||||
| Stable-Core Zero | 243 | |||||
| GPT-Image Zero | 238 | |||||
| Flux-Klein CoT | 231 | |||||
| Stable-Core CoT | 243 | |||||
| GPT-Image CoT | 239 |
| Model | Mean Pullback | Metaphor Consistency | Analogy Appropriateness | Conceptual Integration | ||
| Gemma 4 31B | 5 | 0.05 | 2.8874 | 8.1534 | 8.6399 | 8.6667 |
| Gemma 4 31B | 5 | 0.10 | 2.8100 | 8.0467 | 8.5933 | 8.6000 |
| Gemma 4 31B | 5 | 0.20 | 2.8064 | 8.0467 | 8.5933 | 8.6134 |
| Gemma 4 31B | 10 | 0.05 | 3.1357 | 7.9800 | 8.4800 | 8.6200 |
| Gemma 4 31B | 10 | 0.10 | 3.0188 | 7.9000 | 8.4400 | 8.5667 |
| Gemma 4 31B | 10 | 0.20 | 2.9070 | 7.9400 | 8.4533 | 8.5867 |
| Mean Pullback | % Zero | Spearman vs. CLIP | |
| 0.35 | 2.3144 | 2.4% | 0.0753 |
| 0.40 | 2.2305 | 3.2% | 0.0578 |
| 0.45 | 2.0538 | 7.2% | 0.0690 |
| 0.55 | 1.4883 | 26.4% | -0.0242 |