Organizations: School of Artificial Intelligence, Nanjing University · State Key Laboratory for Novel Software Technology, Nanjing University · Noah’s Ark Lab, Huawei Technologies
Multimodal continual instruction tuning (MCIT) aims to enable multimodal large language models to acquire new capabilities from sequential tasks while preserving previously learned knowledge. Existing methods primarily mitigate catastrophic forgetting by constraining parameter updates or separating task-specific adaptations. However, continual adaptation can also benefit from external knowledge that provides domain-specific information and reusable reasoning patterns for solving diverse instructions. For example, to answer "How many red cubes are to the left of the sphere?", domain knowledge can provide relevant concepts about objects and spatial relations, while reasoning knowledge can specify ordered operations such as object recognition, spatial filtering, and counting. Despite this potential, how to leverage external knowledge for continual adaptation remains largely unexplored in existing MCIT methods. To this end, we propose ReCAP, a retrieval-guided framework that leverages external knowledge to guide capability reuse during continual adaptation. At each continual stage, ReCAP uses external search and an LLM to incrementally build a knowledge base of domain, reasoning, and format knowledge based on the current-stage training data. For each instruction, retrieved domain knowledge guides generation, while retrieved reasoning knowledge selects and orders capability modules to form an instance-specific capability path. As these capability modules are reused across stages, subsequent adaptation can overwrite previously learned parameters. To enable stable cross-stage reuse, ReCAP introduces adaptive subspace recycling, which parameterizes reusable capability modules with shared bases and stage-specific cores, protects historically important directions while recycling residual capacity. Extensive experiments on MCIT benchmarks show that ReCAP achieves SOTA performance.
Figures & tables
Figure 1: An example requiring chart understanding and comparison reasoning.
Figure 2: Overview of ReCAP . (1) At each continual stage, a structured knowledge base is incrementally constructed from the current training instructions with external search and LLM summarization; the current instruction retrieves domain guidance, a reasoning path, and a format schema. (2) Domain guidance augments the model prompt, while the retrieved reasoning path selects and sequentially composes parameterized modules from a shared capability pool. At inference time, instruction, retrieved knowledge, visual, and path-transition cues select a stage-specific core for each activated capability. (3) Adaptive subspace recycling represents recurring capabilities with shared bases and stage-specific cores, protecting high-energy historical directions, recycling low-energy directions, and appending new directions for subsequent adaptation.
Figure 3: Adaptive subspace recycling. Historical capability parameterizations are re-expressed in energy-ordered coordinates without changing their represented functions. High-energy directions are protected, low-energy directions are recycled as trainable capacity, and new directions are appended for subsequent adaptation.
Method
ImageNet-R
ArxivQA
VizWiz
IconQA
CLEVR
Flickr30k
A↑
LoRA-FT ( Hu et al., 2022 )
58.03
77.63
44.39
67.40
61.77
58.22
61.24
O-LoRA ( Wang et al., 2023 )
77.50
78.07
44.50
63.13
64.73
58.16
64.35
MoELoRA ( Chen et al., 2024 )
70.07
77.70
44.69
50.03
54.03
57.34
58.98
ModalPrompt ( Zeng et al., 2025 )
51.07
87.27
48.11
39.23
46.57
42.93
52.53
CL-MoE ( Huai et al., 2025 )
66.33
77.00
44.78
51.87
53.53
57.42
58.49
HiDe-LLaVA ( Guo et al., 2025 )
84.03
90.73
44.43
58.93
41.37
54.25
62.29
Table 1: Comparison with existing methods on UCIT. Task-wise results report final performance Ai,T . Best and second-best values among the main comparison methods are bolded and underlined, respectively; Vanilla RAG variants are shown for reference.
Method
ScienceQA
TextVQA
ImageNet
GQA
VizWiz
Grounding
VQAv2
OCR-VQA
A↑
LoRA-FT ( Hu et al., 2022 )
26.00
25.38
28.51
33.07
26.52
0.10
40.00
52.92
29.06
O-LoRA ( Wang et al., 2023 )
75.40
52.89
71.85
47.30
37.35
7.10
61.85
61.20
51.87
MoELoRA ( Chen et al., 2024 )
62.02
52.05
37.21
53.12
43.32
33.22
57.92
65.75
50.58
ModalPrompt ( Zeng et al., 2025 )
68.42
56.40
41.13
61.11
50.13
36.69
66.90
59.68
55.06
CL-MoE ( Huai et al., 2025 )
73.28
59.94
31.80
60.22
46.98
64.48
67.36
62.39
58.31
SEFE ( Chen et al., 2025 )
75.35
58.66
83.10
54.25
48.85
16.75
65.35
66.25
58.57
Table 2: Comparison with existing methods on CoIN. Task-wise results report final performance Ai,T . Best and second-best values among the main comparison methods are bolded and underlined, respectively; Vanilla RAG variants are shown for reference.
Figure 4: Instance-level capability paths for examples sampled from the UCIT test set. Samples within the same task may require different capability compositions, while shared capabilities can be reused across tasks.
Variant
A↑
ReCAP
70.65
w/o Domain Guidance
70.02
w/o Reasoning-Guided Ordering
70.08
w/o Adaptive Subspace Recycling
49.88
w/o Stage-specific Core Routing
31.50
Table 3: Component ablations of ReCAP on UCIT. Higher is better.
Variant
A↑
ReCAP
70.65
w/o Domain Guidance
70.02
w/o Reasoning-Guided Ordering
70.08
w/o Adaptive Subspace Recycling
49.88
w/o Stage-specific Core Routing
31.50
Table 3: Component ablations of ReCAP on UCIT. Higher is better.
ρ
Protected Rank
A↑
0.50
1.01
67.49
0.70
1.10
66.91
0.90
1.34
69.33
0.99 (default)
2.53
70.67
1.00
8.37
66.52
Table 4: Sensitivity to the historical-energy ratio ρ , with the resulting protected rank.
Table 5: Fixed card-content matching and stage-specific core-selection coefficients used by ReCAP . Signals are listed in the same order as their coefficients.
Setting
α
Domain
Reasoning
Format
Core Selection
Routing Acc. (%)
A↑
Default
0.80
0.58, 0.32, 0.10
0.48, 0.27, 0.10, 0.10, 0.05
0.60, 0.40
0.60, 0.15, 0.20, 0.05
99.304
70.67
Balanced retrieval
0.50
0.33, 0.33, 0.34
0.20, 0.20, 0.20, 0.20, 0.20
0.50, 0.50
default
66.354
67.41
Semantic-emphasized retrieval
0.90
0.65, 0.25, 0.10
0.55, 0.20, 0.10, 0.10, 0.05
default
default
99.141
70.34
Balanced routing
default
default
default
default
0.25, 0.25, 0.25, 0.25
99.761
70.51
Instruction-emphasized routing
default
default
default
default
0.70, 0.10, 0.15, 0.05
98.982
70.22
Appendix
Table 6: Sensitivity to retrieval and stage-specific core-selection coefficients on UCIT. Weight tuples follow the signal order in Table 5 , with α=αdom=αrea . Routing accuracy is computed over all routed capability positions.
Card
Principal stored fields
Forward consumer
In prompt?
Domain
concepts, keywords, knowledge, prompt guidance
generator
prompt guidance only
Reasoning
domains, schemas, ordered operations
capability-path constructor
no
Format
schema, template, validation rule
reasoning retrieval, core router
no
Appendix
Table 7: Knowledge-card fields and their consumers.
δ
Krealmax
Trainable Params (M)
ImgNet-R
ArxivQA
VizWiz
IconQA
CLEVR
Flickr30k
A↑
Δ
0
8
22.176
85.87
93.00
60.70
65.53
59.27
55.97
70.06
−0.61
1 (default)
11
25.340
85.57
92.60
60.56
68.53
60.57
56.09
70.67
0.00
2
14
28.491
85.77
92.80
60.23
69.27
61.87
56.40
71.06
+0.39
3
17
31.879
85.47
92.67
60.66
70.33
62.23
56.53
71.32
+0.65
4
20
35.145
85.80
92.87
60.86
71.10
62.47
56.38
71.58
+0.91
Appendix
Table 8: Sensitivity to rank expansion on UCIT. Krealmax denotes the largest realized active rank, and Trainable Params reports the trainable parameter count under each setting. Δ is relative to the default δ=1 . Best task-wise and average results are bolded.
Method
Trainable Params (M) ↓
Device-h/task ↓
Train samples/s/dev. ↑
SAME ( Xie et al., 2026 )
73.286
7.12
1.391
HiDe-LLaVA ( Guo et al., 2025 )
29.360
3.94
2.510
SEFE ( Chen et al., 2025 )
340.795
8.30
1.192
LoRA-FT ( Hu et al., 2022 )
239.862
3.62
2.735
ReCAP (Ours)
25.340
6.48
1.529
Appendix
Table 9: Trainable-parameter footprint and training efficiency on UCIT. Lower is better for trainable-parameter count and device-hours, whereas higher is better for throughput.
Method
Tokens/s ↑
TTFT (ms) ↓
Samples/s ↑
Peak memory (GB) ↓
Zeroshot
11.4597
198.89
2.0963
14.23
LoRA-FT ( Hu et al., 2022 )
6.2338
246.08
1.4529
15.48
O-LoRA ( Wang et al., 2023 )
8.8487
223.00
2.0312
28.33
MoELoRA ( Chen et al., 2024 )
1.8979
619.01
0.4898
15.44
ModalPrompt ( Zeng et al., 2025 )
9.9822
269.88
1.7951
29.06
CL-MoE ( Huai et al., 2025 )
2.4209
499.94
0.5511
14.83
Appendix
Table 10: Inference efficiency on UCIT. Higher throughput and lower TTFT and peak memory are better.
Multimodal Large Language Models (MLLMs) unify heterogeneous vision-language tasks under a shared generative framework via instruction tuning, yet real-world deployment demands continuous capability expansion, making Multimodal Continual Instruction Tuning (MCIT) essential. Existing methods either update all tasks with a shared parameter set or allocate dedicated modules for each new task. Shared updates force heterogeneous tasks to compete, causing forgetting of learned capabilities. Conversely, isolated expansion prevents interference but severely limits parameter efficiency over long task streams. To address this dilemma, we propose CRAM. Specifically, by isolating task-specific patterns into independent modules, CRAM mitigates catastrophic forgetting across tasks. To further boost parameter efficiency, we utilize adaptive-rank instantiation to identify the capability gap between existing expert capability and new task demands, and dynamically allocate only the necessary parameters. To ensure stable reuse among tasks, centroid-guided routing recognizes and activates existing experts' capabilities, while an orthogonality penalty confines new updates to task-specific directions, preventing re-learning general capability. Extensive experiments across diverse benchmarks consistently demonstrate its superiority over existing methods.
Jun-Tao Tang, Zhen-Hao Xie, Yu-Cheng Shi +1
School of Artificial Intelligence, Nanjing University, China · State Key Laboratory of Novel Software Technology, Nanjing University, China
Multimodal Large Language Models (MLLMs) achieve strong performance through instruction tuning, yet real-world deployment often requires continual capability expansion across sequential tasks. In such scenarios, Multimodal Continual Instruction Tuning (MCIT) aims to acquire new capabilities while limiting catastrophic forgetting. Existing methods mainly follow a module-composition paradigm: they maintain task-level prompts or LoRA experts and dynamically route or aggregate a subset of them at inference. However, samples within the same task can still differ substantially in visual scenes, question intents, and reasoning demands. This motivates instance-level adaptation to individual query-image pairs rather than only selecting or combining task-level modules. To this end, we propose DRAPE (Dynamic Cross-Modal Prompt Generation), a prompt-learning framework that synthesizes continuous instance-specific soft prompts for MCIT. Instead of selecting prompts from a fixed pool, DRAPE derives prompt queries from the textual instruction and cross-attends to visual patch features, producing query-image conditioned prompts that are prepended to the frozen LLM. To mitigate forgetting during sequential updates, DRAPE applies null-space gradient projection to the shared projector and uses CLIP-based prototype routing for task-label-free generator selection at inference. Extensive experiments on MCIT benchmarks show that DRAPE achieves state-of-the-art performance among representative prompt-based and LoRA-based continual-learning baselines.
Tao Hu, Da-Wei Zhou
School of Artificial Intelligence, Nanjing University · State Key Laboratory for Novel Software Technology, Nanjing University
Multimodal Large Language Models (MLLMs) rely on a projector to align visual representations with the language embedding space, making it central to cross-modal understanding. In Multimodal Continual Instruction Tuning (MCIT), however, shifting visual distributions and evolving instruction semantics cause this shared projector to drift, leading to projector-level forgetting, an issue largely overlooked by methods that focus primarily on the LLM backbone. We introduce Progressive Multimodal Alignment (PMA), a framework that enables the projector to adapt continually while preserving previously learned alignment. PMA detects multimodal distribution shifts via a lightweight representation descriptor and progressively expands projector experts only when needed. An expandable router integrates expert outputs based on multimodal features, while the original pretrained projector is retained as a stable alignment anchor. This progressive mechanism balances stability and plasticity with sub-linear parameter growth and serves as a method-agnostic add-on to existing MCIT approaches. Extensive experiments on two recent MCIT benchmarks demonstrate that mitigating projector-level forgetting yields consistent gains over prior state-of-the-art methods when combined with PMA. Moreover, PMA scales across diverse MLLM backbones, demonstrating robust and broadly applicable MCIT performance.
Duzhen Zhang, Yahan Yu, Qiaoyi Su +2
Mohamed bin Zayed University of Artificial Intelligence · Center for Excellence in Brain Science and Intelligence Technology, Chinese Academy of Sciences · Kyoto University +2