Rethinking Multi-Image Re-Representation in Multi-Image Understanding
Organizations: LMU Munich · MCML
Abstract
Multi-image understanding requires MLLMs not only to recognise the content of individual images, but also to organise visual evidence distributed across them. We study this problem through multi-image re-representation, viewing prompted Chain-of-Thought reasoning and agentic visual tool use as different ways of re-organising visual evidence during reasoning. We introduce Mosaic, a general-purpose multi-image visual harness that enables an MLLM to actively construct visual intermediates with ten composable image operations. We compare five re-representation settings on existing multi-image benchmarks and on MosaicBench, a new grounding-focused benchmark for fine-grained multi-image understanding. Our experiments show that the relative benefits of textual and visual re-representation are strongly task-dependent. Visual re-representation is particularly effective for tasks requiring precise visual evidence, including hypothesis testing, precision comparison, and orientation-sensitive reasoning, while tasks dominated by higher-level semantic content show smaller or less consistent gains. Building on this finding, we train MosaicAgent-8B to use Mosaic with reinforcement learning using only accuracy and format rewards. Without demonstration trajectories or rewards for specific tool-use, the agent learns to compose visual operations over multiple steps and exhibits diverse problem-solving patterns unpromptedly. Code and data will be released at https://github.com/gengyuanmax/Mosaic.
Figures & tables
| Model | MosaicBench | M4Bench | Mantis | BLINK | MMIU | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Resol. | Orient. | Prec.Comp. | Hyp.Test | Ctx.Int. | Spat.Ref. | Overall | D.Diff | S.Comp | I.Comp | Overall | Overall | Overall | Overall | |
| Open-weight baselines | ||||||||||||||
| InternVL3.5-8B ( Wang et al., 2025b ) | .438 | .392 | .390 | .167 | .463 | .292 | .363 | .218 | .649 | .228 | .373 | .696 | .581 | .524 |
| LLaVA-OV-1.5-8B ( An et al., 2025 ) | .338 | .375 | .340 | .267 | .400 | .233 | .325 | .000 | .601 | .254 | .315 | .567 | .460 | .417 |
| GLM-4.6V-Flash ( GLM-V Team, 2025 ) | .625 | .475 | .690 | .300 | .600 | .342 | .505 | .606 | .740 | .197 | .536 | .737 | .695 | .627 |
| MiniCPM-V-4.5 ( Yao et al., 2025 ) | .500 | .475 | .550 | .267 | .500 | .425 | .463 | .349 | .664 | .254 | .466 | .728 | .612 | .552 |
| Training | MosaicBench | Mantis | BLINK | M4Bench | MMIU |
|---|---|---|---|---|---|
| Mosaic w/o RL | 45.2 1.44 | 81.9 1.00 | 64.4 0.40 | 51.3 1.20 | 61.0 |
| RL w/o Mosaic | 53.0 0.54 | 80.4 1.47 | 58.2 0.60 | 51.6 0.80 | 56.8 |
| RL w/ Mosaic | 63.3 1.10 | 81.6 1.20 | 62.8 1.00 | 61.8 1.40 | 57.2 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| ID | Task | Source | # Images |
|---|---|---|---|
| Resolution | |||
| 1.1 | Small-target cross-frame tracking | VisDrone2019-MOT | 4 |
| 1.3 | Distant sign reading | TT100K | 1 |
| 1.5 | Pathology zoom | CAMELYON16 | 1 |
| 1.6 | Micro-scratch detection | MVTec AD | 1 |
| 1.7 | Dense small-object counting | CARPK | 1 |
| Source Dataset | Annotations | Task IDs |
|---|---|---|
| VisDrone2019-MOT | Aerial video frames and object-track annotations | 1.1, 4.2, 6.1 |
| TT100K | Street images and traffic-sign annotations | 1.3, 2.5 |
| CAMELYON16 | Whole-slide images and lesion annotations; slide backgrounds for contrast tasks | 1.5, 5.5, 6.6 |
| MVTec AD | Normal images, annotated defect images, and defect patches | 1.6, 2.3, 2.8, 3.10, 3.11, 4.3, 5.6, 6.7 |
| CARPK | Aerial parking-lot images and car annotations | 1.7, 6.8, 6.9 |
| PubLayNet | Document-page images | 2.10 |
| Field | Content |
|---|---|
| sample_id | Sample identifier. |
| prompt | Task question, input layout, and task-specific instructions. |
| images | Input-image references in the order used by the prompt. |
| task_type | Task identifier corresponding to Table 3 . |
| target | Ground-truth answer used for scoring. |
| metadata | Task-specific construction and audit information. |
| Probe | Scope | Decision rule |
|---|---|---|
| Centroid / outlier | Matrix-based tasks | Select the transformation nearest to or furthest from the mean of the supplied transformations. |
| Pair member / singleton | Text-orientation recovery (2.10) | Select an angle from the pair separated by , or the angle outside that pair. |
| Central- / lowest- | Drivable-zone point query (6.4) | Select the most horizontally central point or the lowest point in the image. |
| Tool successes out of eight runs | ||||||||||
| 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | Total | |
| Constructed examples | 3116 | 1647 | 1268 | 995 | 769 | 756 | 715 | 930 | 2604 | 12,800 |
| Incorrect without tools | 2603 | 1429 | 1026 | 735 | 546 | 449 | 395 | 358 | 669 | 8,210 |
| the training suite | – | – | 1026 | 735 | 546 | 449 | 395 | 358 | – | 3,509 |
| Held-out hard subset | 2603 | 1429 | – | – | – | – | – | – | – | 4,032 |
| Remaining examples | 513 | 218 | 242 | 260 | 223 | 307 | 320 | 572 | 2604 | 5,259 |
| Available Visual Tools | MosaicBench | Mantis | BLINK | M4Bench | MMIU |
|---|---|---|---|---|---|
| Crop only | .539 | .806 | .624 | .588 | .583 |
| Full toolset | .633 .011 | .816 .012 | .628 .010 | .618 .014 | .572 |
| Model | Setting | MosaicBench | Mantis | BLINK | M4Bench | MMIU |
|---|---|---|---|---|---|---|
| Qwen3-VL-8B | T- | .452 .043 | .819 .010 | .644 .004 | .513 .012 | .610 |
| V- | .472 .015 | .797 .016 | .636 .002 | .518 .003 | .566 | |
| MosaicAgent-8B | T- | .460 .012 | .828 .004 | .637 .004 | .513 .014 | .594 |
| V- | .633 .011 | .816 .012 | .628 .010 | .618 .014 | .572 |