cs.CVApr 15, 2026

FiRe: Fine-grained Multimodal Reasoning for Enhanced Image Generation

Authors: Yongjin Kim, Yoonjin Oh, Yerin Kim, Hyomin Kim, Jeeyoung Yun, Yujung Heo, Minjun Kim, Sungwoong Kim

Organizations: Department of Artificial Intelligence Korea University Seoul, Republic of Korea · KT Corporation Seoul, Republic of Korea

Abstract

With the rapid progress of Multimodal Large Language Models (MLLMs), unified MLLMs that jointly perform image understanding and generation have advanced significantly. However, despite the inherent reasoning capabilities of unified MLLMs for self-reflection and self-refinement, their use in text-to-image generation remains largely underexplored. Meanwhile, existing multimodal reasoning-based image generation methods mostly rely on prompt augmentation or holistic image-text alignment judgments, without fine-grained reflection and refinement of detailed prompt attributes, leading to limited fine-grained control. To address this limitation, we propose FiRe, a Fine-grained Multimodal Reasoning method for enhanced image generation by MLLM. In specific, FiRe performs a fine-grained multi-step reasoning by first decomposing the prompt into key visual requirements and then self-judging their satisfaction in the generated image, followed by localized refinement according to self-generated precise feedback. In addition, to further strengthen the MLLM's multimodal reasoning ability, we introduce FiRe-GRPO, a reinforcement learning method tailored to FiRe. Since standard Group Relative Policy Optimization (GRPO) suffers from sparse, outcome-based rewards in multi-step reasoning, we formulate our reasoning process as a step-level decision-making problem, design step-specific rewards, and compute step-level advantages for granular credit assignment within GRPO. Extensive experiments demonstrate that FiRe consistently outperforms competitive text-to-image baselines, including existing reasoning-based methods, with particularly substantial gains on compositional text-to-image benchmarks. Our project page is available at https://ku-agi.github.io/FiRe/

Explore similar work

CardsList
  1. Large Language Models are Universal Reasoners for Visual Generation

    May 5, 2026Sucheng Ren, Chen Chen, Zhenbang Wang +5Visual GenerationFeature Alignment

  2. Benchmarking and Evolving Reason-Reflect-Rectify for Reflective Visual Generation

    May 19, 2026Junjie Wang, Xinghua Lou, Jason Li +8Visual GenerationMultimodal Generation