ReMAP: Restoring the Perceptual Cycle with Reasoning-Time Latent Visual Memory
Organizations: University of Electronic Science and Technology of China · University of Michigan, Ann Arbor · Nanyang Technological University · Peking University · Mohamed bin Zayed University of Artificial Intelligence · New York University · Harvard University · Massachusetts Institute of Technology
Abstract
As multimodal large language models (MLLMs) reason for longer, attention to the initial visual input diminishes, weakening visual grounding. Visual memory reintroduces visual evidence during reasoning. We conduct a controlled analysis of visual memory along three axes: curation, organization, and access. We find that local evidence benefits from global context, compact latent representations balance accuracy and visual-context cost, and the utility of memory access depends on the reasoning state. Guided by these findings, we propose ReMAP (Reasoning-Time Memory-Augmented Perception), which couples two complementary latent memories: a static, question-conditioned Global memory that preserves scene and cross-image context, and a dynamic Local memory that uses this context as an anchor while selecting and re-encoding region-level evidence according to the current reasoning state. Both memories return compact latent tokens inserted into the reasoning sequence, and a reinforcement-learning access policy trained with branched rollouts decides when to continue reasoning or invoke Global or Local memory. On ten benchmark families, ReMAP outperforms prior visual-memory methods on all four multi-image benchmarks, exceeding the strongest prior results on MuirBench and MIMIC by 8.38 and 14.84 percentage points. Across four backbone families, enabling memory access improves over the same trained model with memory disabled, and on shared V*Bench, CV-Bench-2D, and MuirBench questions ReMAP reduces the visual tokens entering the reasoning sequence by 51.0-76.8% relative to the native-resolution backbone. Further analyses show that Global and Local memory form distinct yet complementary latent representations. Together, these components restore the perceptual cycle by letting the reasoning state trigger targeted visual retrieval, with the retrieved evidence guiding subsequent reasoning.
Figures & tables
| Single-image evidence acquisition | Multi-image relational reasoning | Held-out generalization | ||||||||||
| Method | V ⋆ | CV-2D | CV-3D | HR-4K | HR-8K | BLINK | Muir | MMSI | MIMIC Avg. | RWQA | MathVista | MMVet |
| Native | 76.44 | 74.41 | 76.33 | 72.88 | 67.75 | 56.62 | 57.85 | 28.10 | 45.17 | 67.97 | 59.90 | 46.06 |
| Coarse | 67.54 | 73.10 | 75.42 | 61.13 | 54.62 | 57.79 | 58.15 | 28.40 | 43.56 | 67.19 | 59.50 | 47.06 |
| Pixel reinspection | ||||||||||||
| Q-Zoom | 79.06 | 76.35 | 74.08 | 73.00 | 66.63 | 53.97 | 50.38 | 25.40 | 42.29 | 70.85 | 54.60 | 47.43 |
| DeepEyes | 84.29 | 78.61 | 79.42 | 74.63 | 68.75 | 49.24 | 48.50 | 24.10 | 39.27 | 69.67 | 62.00 | 51.01 |
| Method | V ⋆ | CV-3D | Muir | BLINK | MM-Vet |
|---|---|---|---|---|---|
| Base | 67.54 | 75.42 | 58.15 | 57.81 | 47.06 |
| SFT | 58.64 | 75.58 | 55.73 | 54.13 | 55.05 |
| Stage I sft | 76.44 | 83.33 | 57.81 | 58.92 | 48.35 |
| Memory curation | |||||
| w/o Memory | 79.06 | 80.00 | 61.00 | 59.02 | 57.80 |
| w/ Global memory | 79.58 | 83.00 | 64.00 | 59.50 | 60.55 |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Pixel / Region Reacquisition | Compact Visual Memory | State- Conditioned Access | Cross-Image Context | Global / Local Split | Invocation vs. Continuation |
| DeepEyes | ✓ | ✗ | – | – | ✗ | ✗ |
| Pixel Reasoner | ✓ | ✗ | – | – | ✗ | ✗ |
| VLM-R 3 | ✓ | ✗ | – | – | ✗ | ✗ |
| CMMCoT | ✓ | ✓ | – | ✓ | ✗ | ✗ |
| Qwen Look Again | ✗ | ✓ | – | – | ✗ | ✗ |
| ICoT / MINT-CoT | ✗ | ✓ | – | – | ✗ | ✗ |
| Stage | Parameter group or setting | Value |
|---|---|---|
| I: Global formation | Image and Relation Weavers | |
| Output projection | ||
| I: Local formation | Source Selector, Twig, and Local Weaver | |
| I: interface warm-up | Marker and request rows | |
| II: access learning | Decoder LoRA | |
| Decoder LoRA rank / | / |
| Source | Content and use |
|---|---|
| Visual-CoT (GQA) ( Shao et al., 2024a ) | Object, attribute, and spatial-relation questions with region annotations. |
| COCO train2017 ( Lin et al., 2014 ) | Multi-image questions built from instance annotations. |
| Mantis-Instruct ( Jiang et al., 2024 ) : ChartQA, Spot-the-Diff, ImageCoDe, NLVR2, and Birds-to-Words | Chart reading and arithmetic, image comparison, description-to-image matching, and cross-image relational judgments. |
| Q-Zoom’s VStar-COCO data ( Shi et al., 2026 ) | Spatial-relation questions on COCO train2017 images. |
| ScanQA via M4-Instruct ( Li et al., 2024 ) | Multi-view scene question answering, restricted to examples from the original ScanQA training split. |
| Evaluation unit | Examples | Aggregated metric |
| Single-image perception | ||
| V*Bench, released test | 191 | Micro accuracy |
| CV-Bench, test-2D | 1,438 | Mean source accuracy |
| CV-Bench, test-3D | 1,200 | Micro accuracy |
| HR-Bench, 4K | 800 | Mean category accuracy |
| HR-Bench, 8K | 800 | Mean category accuracy |
| Method | Common | Counting | Odd-One-Out | Listing | Mean |
|---|---|---|---|---|---|
| Native | 62.60 | 44.98 | 60.40 | 12.69 | 45.17 |
| Coarse | 61.10 | 38.42 | 54.70 | 20.02 | 43.56 |
| Q-Zoom † | 51.90 | 37.44 | 54.80 | 25.00 | 42.29 |
| DeepEyes † | 52.60 | 40.38 | 53.20 | 10.90 | 39.27 |
| Pixel Reasoner | 64.50 | 43.34 | 30.00 | 41.46 | 44.83 |
| ReGround | 53.60 | 36.76 | 13.00 | 8.66 | 28.01 |
| Cohort | Eligible / selected | First last bin | Paired change |
|---|---|---|---|
| V*Bench | [ ] pp | ||
| BLINK | [ ] pp | ||
| MuirBench | [ ] pp | ||
| Pooled | [ ] pp |
| Cohort | Contrast | Estimate | 95% CI |
|---|---|---|---|
| V*Bench ( ) | late–early correct-content utility | ||
| V*Bench ( ) | correct–random-wrong RoI | NS | |
| CV-Bench-3D ( ) | late–early current-scene utility | ||
| CV-Bench-3D ( ) | current–matched-wrong scene | ||
| CV-Bench-3D ( ) | current–shuffled scene | NS |
| Contrast | Estimate | 95% CI |
|---|---|---|
| Role change appearance control | ||
| Late early role sensitivity | NS | |
| Late early fixation rate | NS | |
| Oracle no context | ||
| Oracle blank block | ||
| Oracle zero block |
| Benchmark | Method | Initial | Added | Total |
|---|---|---|---|---|
| V*Bench ( ) | Native | 4,417.9 | 0.0 | 4,417.9 |
| Q-Zoom | 1,008.0 | 331.8 | 1,339.8 | |
| DeepEyes | 4,417.9 | 35.9 | 4,453.8 | |
| VisMem | 4,417.9 | 0.0 | 4,417.9 | |
| ReMAP | 1,008.0 | 16.0 | 1,024.0 | |
| CV-Bench-2D ( ) | Native | 906.0 | 0.0 | 906.0 |