CoDe-LoRA: Mitigating the Orthogonality Dilemma in Continual Learning of LLMs via Knowledge Consolidation and Decoupling
Organizations: Beijing University of Posts and Telecommunications · National University of Singapore
Abstract
Continual learning (CL) is essential for Large Language Models (LLMs) to sequentially adapt to evolving tasks. To mitigate catastrophic forgetting, recent advances implement low-rank adaptation with orthogonal projections (e.g., O-LoRA) to isolate task parameters. However, we reveal that such strict geometric constraints trigger an "Orthogonality Dilemma": rigid parameter isolation impedes the transfer and accumulation of shared representations across semantically related tasks. In this work, we propose a new replay-free method, called Consolidation and Decoupling LoRA (CoDe-LoRA), for CL of LLMs. CoDe-LoRA disentangles the learning process into Consolidating Universal Knowledge and Decoupling Task-Specific Knowledge. To achieve this, CoDe-LoRA leverages an adaptive null space projection mechanism and semantic routing to balance knowledge accumulation with task-specific adaptation. Experimental results across four backbones and three CL benchmarks show that CoDe-LoRA achieves the best average accuracy. Our code is available at https://github.com/Estrellajer/CoDe-LoRA.
Figures & tables
| Method | Standard CL Benchmark | Large Number of Tasks | |||||||
| Order-1 | Order-2 | Order-3 | Avg | Order-4 | Order-5 | Order-6 | Avg | ||
| Ref. | Per-LoRA (Oracle) | 81.4 | 81.3 | 81.4 | 81.4 | 77.6 | 77.4 | 77.5 | 77.5 |
| MTL | 80.0 | 80.0 | 80.0 | 80.0 | 76.5 | 76.5 | 76.5 | 76.5 | |
| Seq. FT | SeqFT | 18.9 | 24.9 | 41.7 | 28.5 | 7.4 | 7.4 | 7.5 | 7.4 |
| SeqLoRA | 44.6 | 32.7 | 53.7 | 43.7 | 0.6 | 1.9 | 1.6 | 1.4 | |
| IncLoRA | 66.0 | 64.9 | 68.3 | 66.4 | 63.3 | 58.5 | 61.7 | 61.2 | |
| Method | Train (s) | Infer (s) | Params |
| O-LoRA | 108.8 | 30.1 | 9.4M (1.28%) |
| N-LoRA | 98.0 | 29.9 | 9.4M (1.28%) |
| MoLE-CIE | 861.0 | 142.7 | 32.4M (4.40%) |
| CoDe-LoRA | 98.4 | 43.1 | 9.4M (1.28%) |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Notation | Description |
| The -th task in the sequence | |
| Dataset associated with task | |
| Parameters of the frozen backbone | |
| Parameters of consolidation branch | |
| Parameters of -th task expert | |
| Accumulated shared weights up to task |
| Method | Avg Accuracy | Shared Branch | |
| No Scaling | |||
| Linear Avg | |||
| CoDe-LoRA |
| Retained rank | Co-LoRA alone | CoDe-LoRA |
| 1 | 66.84 | 79.13 |
| 2 | 72.11 | 79.13 |
| 4 | 75.86 | 79.13 |
| 8 | 76.67 | 79.13 |
| Standard (Order 1) | Long (Order 4) | |||
| Router | Misrouted | Loss | Misrouted | Loss |
| Cosine prototype | 2.03% | 1.43 | 10.07% | 4.74 |
| LDA | 0.15% | 0.10 | 1.24% | 0.20 |
| Scheme family | Outcome |
| Confidence-threshold fallback | best threshold on development is “never use Co” |
| Per-expert branch choice | routed expert |
| Co as an extra expert | no development-qualified candidate |
| Output selectors (margin, entropy, agreement) | below calibrated De, or kept De |
| 1480-feature selectors | none passed development; 26 rescues vs. 26 breaks on test |
| Closed-form combinations (PoE, averaging, Dawid–Skene) | do not solve the selection problem |
| Construction | Co alone | Oracle |
| Trained, SVD merge (main) | 76.67 / 66.33 | 83.79 / 84.61 |
| + null space, paper scaling | 75.61 / 63.78 | 84.53 / 86.00 |
| + distillation after SVD | 74.77 / 66.34 | 83.66 / 84.51 |
| + left-space projection | 76.16 / 65.67 | 83.74 / 84.08 |
| + two-sided projection | 74.88 / 63.91 | 83.14 / 84.60 |
| Isolated experts, SVD merge | 41.87 / 11.06 | 81.84 / 73.99 |
| Method | Task | Train (s) | Infer (s) |
| O-LoRA | DBpedia | 141 | 25 |
| Amazon | 65 | 27 | |
| Yahoo | 186 | 43 | |
| AG News | 44 | 117 | |
| N-LoRA | DBpedia | 128 | 26 |
| Amazon | 60 | 27 |
| Method | Train (min) | Infer (min) | Params |
| O-LoRA | 13.3 | 9.4 | 35.4M (4.80%) |
| N-LoRA | 12.0 | 9.4 | 35.4M (4.80%) |
| MoLE-CIE | 140.3 | 49.5 | 60.0M (8.14%) |
| CoDe-LoRA | 12.5 | 13.4 | 35.4M (4.80%) |
| No. | Dataset name | Category | Task | Domain | Metric |
| 1 | Yelp | CL Benchmark | sentiment analysis | Yelp reviews | accuracy |
| 2 | Amazon | CL Benchmark | sentiment analysis | Amazon reviews | accuracy |
| 3 | DBpedia | CL Benchmark | topic classification | Wikipedia | accuracy |
| 4 | Yahoo | CL Benchmark | topic classification | Yahoo Q&A | accuracy |
| 5 | AG News | CL Benchmark | topic classification | news | accuracy |
| 6 | MNLI | GLUE | NLI | various | accuracy |
| Task | Prompts |
| NLI | What is the logical relationship between the “sentence 1” and the “sentence 2”? Choose one from the option. |
| QQP | Whether the “first sentence” and the “second sentence” have the same meaning? Choose one from the option. |
| SC | What is the sentiment of the following paragraph? Choose one from the option. |
| TC | What is the topic of the following paragraph? Choose one from the option. |
| BoolQA | According to the following passage, is the question true or false? Choose one from the option. |
| MultiRC | According to the following passage and question, is the candidate answer true or false? Choose one from the option. |
| Order | Setting | Task Sequence |
| 1 | T5, Llama 2-7B | dbpedia amazon yahoo ag |
| 2 | T5, Llama 2-7B | dbpedia amazon ag yahoo |
| 3 | T5, Llama 2-7B | yahoo amazon ag dbpedia |
| 4 | T5 | mnli cb wic copa qqp boolqa rte imdb yelp amazon sst-2 dbpedia ag multirc yahoo |
| 5 | T5 | multirc boolqa wic mnli cb copa qqp rte imdb sst-2 dbpedia ag yelp amazon yahoo |
| 6 | T5 | yelp amazon mnli cb copa qqp rte imdb sst-2 dbpedia ag yahoo multirc boolqa wic |