Organizations: Tsinghua Shenzhen International Graduate School, Tsinghua University · School of Vehicle and Mobility, Tsinghua University · SZ DJI Technology Co., Ltd. · Tencent Holdings Limited
Task-specific fine-tuning can rewrite a language model's answers beyond the training task, complicating updates that must preserve existing behavior. We introduce ATLAS, which turns retained-domain representations into an input-dependent rule for task adaptation. An activation atlas supplies local reference centers and directional filters to a shared low-rank residual. Target supervision learns the residual, while retained geometry shapes its action throughout training and inference. On Qwen3-8B, ATLAS achieves lower mean retained-output Kullback-Leibler (KL) divergence than all seven published baselines at shared coding-performance requirements, with consistent advantages across multiple training seeds. Structural comparisons identify the contributions of retained reference states and directional conditioning, and answer-level analyses show fewer rewritten mathematical answers and more stable commonsense choices. Experiments spanning five backbones and two retained domains further demonstrate coding gains with reduced retained-output movement. With compact storage and modest decoding overhead, ATLAS provides a practical mechanism for acquiring specialized skills while maintaining continuity in existing responses.
Figures & tables
Figure 1: Learning a task with less answer rewriting. Both adaptation paths start from the same model. Gray marks and orange waves schematically denote unchanged and changed answers. Retained inputs illustrate domains evaluated separately. Atlas charts summarize activation neighborhoods with reference centers and local directions. The same token state h drives routing and enters a shared residual conditioned on the selected geometry.
Method
Mean KL ( ×10−3 ) ↓
ATLAS reduction (%)
Activation and directional constraints
CorDA-KPM
1.402
17.60
LoRA-Null
1.360
15.09
OPLoRA ( k=16 )
1.329
13.07
OPLoRA ( k=128 )
1.312
11.97
CLoRA
1.319
12.42
Table 1: Lower retained KL than seven published baselines. Qwen3-8B, GSM8K, three seeds, and 148–160 solved MBPP+ problems. Mean KL averages requirements within each seed, then summarizes seeds geometrically. All methods have 131,072 trainable parameters except TopLoRA with 135,179.
Residual structure
Mean KL ( ×10−3 ) ↓
ATLAS reduction (%)
Reference states and distances
Residual Plain
1.9718
41.30
Centered residual
1.2138
4.64
Distance-scaled residual
1.1849
2.31
Geometric conditioning
PCA-derived scaling
1.1656
0.70
Table 2: Residual structures at common coding requirements. Qwen3-8B/GSM8K, three seeds, and 151–160 solved problems. Mean KL follows Table 1 ’s aggregation; reductions compare Atlas with each row. Appendix D reports individual seeds.
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Use
Examples
Relation to other uses
MBPP
Target training
120
Training partition
MBPP
Validation NLL
43
Validation partition
MBPP
Code generation
257
Test partition
MBPP+
Pass@1
224
Fixed intersection of test problems
GSM8K
Atlas fitting / Replay
2,048
Same training examples
GSM8K
Retained KL
512
Subset of behavior test
Appendix
Table 3: Data roles. Atlas fitting and Replay share their training examples. Retained-KL subsets belong to the corresponding behavior evaluation set. MBPP training, validation, and test use separate source partitions.
Backbone identifier
Layer
Rank
Batch
Learning rate
Max. tokens
Qwen/Qwen3-4B-Base
17
16
2×8
2×10−4
1,024
Qwen/Qwen3-8B-Base
17
16
1×16
2×10−4
1,024
Qwen/Qwen3.5-9B-Base
15
16
1×16
2×10−4
1,024
Qwen/Qwen3-14B-Base
19
16
1×16
2×10−4
1,024
ZhipuAI/glm-4-9b-hf
19
16
1×16
4×10−4
1,024
Appendix
Table 4: Backbone identifiers and ATLAS training settings. Layer indices are zero-based; batch denotes target microbatch size multiplied by gradient-accumulation steps. The listed learning rates apply to all three training seeds.
Method
Rank
Mechanism parameters
Learning rate
LoRA-Null
16
Lowest 16 activation-covariance eigenvectors for initialization
2×10−4
OPLoRA
16
Weight-SVD complement projections, k=16
2×10−4
OPLoRA
16
Weight-SVD complement projections, k=128
2×10−4
CorDA-KPM
16
Smallest 16 context-oriented singular components
4×10−4
TopLoRA
11
α=16 ; token scale exp(RMSNorm(Cx)) ; dropout 0
4×10−4
STM
16
Retain response tokens with starting-model perplexity ≤2.5
1×10−4
Appendix
Table 5: Qwen3-8B weight-adaptation settings. Every method updates the same attention-output matrix. OPLoRA’s protected subspace width k is separate from its trainable rank. The learning rates apply to all three training seeds.
MBPP+ (%)
ATLAS / comparator KL ↓
Model
Retained
Frozen
Plain
ATLAS
Plain
Replay
Output-KL
Qwen3-4B
GSM8K
2.23
35.71
14.73
0.441
0.398
0.400
Qwen3-8B
GSM8K
50.89
72.17
71.43
0.525
0.570
0.688
Qwen3.5-9B
GSM8K
62.50
63.54
65.62
0.202
0.218
0.309
Qwen3-14B
GSM8K
65.18
71.13
72.62
0.579
0.363
0.618
GLM-4-9B
GSM8K
36.61
65.77
70.54
0.312
0.194
0.581
Appendix
Table 6: Target learning and retained-output movement across model scales and domains. Target scores average three validation-selected checkpoints. KL ratios compare complete trajectories over pairwise shared integer target requirements. Ratios below one favor Atlas . All five backbones and both retained domains are included.
Backbone / retained domain
Comparator
7302026
7302027
7302028
Qwen3-4B
Plain
0.444 [12,35]
0.465 [11,37]
0.416 [11,27]
Qwen3-4B
Replay
0.343 [12,35]
0.429 [11,37]
0.428 [10,27]
Qwen3-4B
Output-KL
0.458 [12,35]
0.378 [8,37]
0.370 [13,27]
Qwen3-8B
Plain
0.578 [150,161]
0.526 [144,161]
0.477 [151,163]
Qwen3-8B
Replay
0.571 [149,161]
0.581 [148,161]
0.558 [151,162]
Qwen3-8B
Output-KL
0.621 [147,161]
0.686 [145,161]
0.765 [148,163]
Appendix
Table 7: Primary target-score comparisons, by training seed. Each cell gives the ratio of mean retained KL (ATLAS/comparator), followed by the inclusive integer MBPP+ range [kmin,kmax] . Every value uses observed checkpoints meeting each target threshold. GSM8K is retained unless CSQA is named.
Backbone / retained domain
7302026
7302027
7302028
Qwen3-4B
0.8718
0.8797
0.9278
Qwen3-8B
0.9721
0.9746
0.9814
Qwen3.5-9B
0.9553
0.9911
0.9593
Qwen3-14B
0.9727
0.9797
1.0002
GLM-4-9B
2.9425
3.5886
4.4319
Qwen3-8B / CSQA
0.9751
0.9595
0.9566
Appendix
Table 8: ATLAS/Global-PCA mean-KL ratios on each pair’s shared integer target range. Values below one favor the local atlas. GSM8K is retained unless CSQA is named.
Method
Selected learning rate
CorDA-KPM
4×10−4
TopLoRA
4×10−4
CLoRA
4×10−4
STM
10−4
TALR
2×10−4
Appendix
Table 9: Learning-rate selection for the published-method comparison. Each listed method evaluates 10−4 , 2×10−4 , and 4×10−4 on seed 7302026. Minimum validation NLL over epochs 1–5 selects the rate used by all three seeds.
Comparator
7302026
7302027
7302028
Geometric mean
CLoRA
0.8893
0.8747
0.8636
0.8758
CorDA-KPM
0.7780
0.7490
0.9600
0.8240
LoRA-Null
0.8466
0.8566
0.8442
0.8491
OPLoRA ( k=128 )
0.8887
0.8768
0.8755
0.8803
OPLoRA ( k=16 )
0.8789
0.8623
0.8668
0.8693
STM
0.9202
0.8843
0.8912
0.8984
Appendix
Table 10: External methods at the same 13 target thresholds, k∈{148,…,160} out of 224 MBPP+ problems. Entries are ATLAS/comparator ratios of mean retained KL. The last column geometrically averages the three seed ratios.
Method
MBPP+ (%)
GSM8K (%)
KL ( ×10−3 )
Reduction (%)
ATLAS
71.43
86.38
1.2876
—
CorDA-KPM
64.58
86.58
10.7280
88.00
LoRA-Null
72.02
86.28
4.1632
69.07
OPLoRA ( k=16 )
71.13
86.58
4.0335
68.08
OPLoRA ( k=128 )
71.43
86.48
3.7502
65.66
CLoRA
69.49
86.50
5.0130
74.31
Appendix
Table 11: Independent validation selection on Qwen3-8B. Each trajectory contributes its checkpoint with the lowest target-validation NLL. MBPP+ and GSM8K are arithmetic means over three training seeds; retained KL is a geometric mean. KL reduction is one minus the geometric mean of paired ATLAS-to-comparator ratios, expressed as a percentage.
Structure
Seed
Observed range
Mean KL ×103
KL ratio
Residual Plain
7302026
150–162
2.0773
0.55920
7302027
144–162
2.0341
0.56939
7302028
151–163
1.8143
0.63536
Centered residual
7302026
144–163
1.2251
0.94819
7302027
147–163
1.2092
0.95781
7302028
149–162
1.2073
0.95481
Appendix
Table 12: Residual structures at the same ten integer target requirements, 151–160 solved MBPP+ problems out of 224. Each KL value averages the best achieved KL over these requirements within one training seed. Ratios divide the ATLAS value by the corresponding method value. Observed ranges describe all five checkpoints in each trajectory.
Setting
Scenarios
Centered
Distance-scaled
Measured counts and requirements
1
4.64
2.31
Omit one target requirement
10
4.44–4.79
2.06–2.46
One checkpoint count ±1
60
4.33–5.08
2.04–2.50
ATLAS counts −1 ; comparator counts +1
1
3.62
1.10
Appendix
Table 13: ATLAS retained-KL reductions (%) under target-score sensitivity. Ranges span the stated scenarios per comparator; every scenario favors ATLAS in all three seeds.
Variant
Emitted residual Δ(h)
Plain
Btanh(Ah)
Centered residual
Btanh(Ax)
Distance-scaled residual
tanh(∥x∥2/sd)Btanh(Ax)
PCA-derived scaling
ρ(h)Btanh(Ax)
Input filter
ρ(h)Btanh(AN(h)x)
Output filter
ρ(h)N(h)Btanh(Ax)
Appendix
Table 14: Residual equations. PCA-based scaled variants share ρ(h) ; distance scaling uses tanh(∥x∥2/sd) . The matrices A,B are trainable.
Backbone
Seed
Epochs
ATLAS KL
Rescaled Plain KL
Reduction (%)
Qwen3.5-9B
7302026
4/2
0.291723
0.310847
6.152
Qwen3.5-9B
7302027
3/5
0.270298
0.279181
3.182
Qwen3.5-9B
7302028
2/2
0.263262
0.266117
1.073
Qwen3-14B
7302026
5/4
0.776070
0.890099
12.811
Qwen3-14B
7302027
2/1
0.695327
0.728186
4.512
Qwen3-14B
7302028
2/4
0.703477
0.738552
4.749
Appendix
Table 15: Twelve effective-norm comparisons on GSM8K. The Plain residual is rescaled token by token to match the effective residual norm of its paired ATLAS checkpoint. Epochs identify ATLAS/Plain. KL columns use units of 10−3 ; positive reduction favors ATLAS.
Comparator
Seed
eA
eB
aA
aB
aA−aB
Qwen3.5-9B, retained GSM8K
Plain
7302026
4
2
147
147
0
Replay
7302026
2
1
146
146
0
Output-KL
7302026
3
1
147
147
0
Plain
7302027
3
3
145
145
0
Replay
7302027
2
1
144
144
0
Appendix
Table 16: GSM8K behavior pairings. Epochs and MBPP+ counts identify both checkpoints in each comparison. The GLM rows cover the Plain, Replay, and Output-KL comparisons summarized in Table 19 .
Comparator
Seed
eA
eB
aA
aB
aA−aB
Qwen3-8B, retained CSQA
Plain
7302026
2
5
160
160
0
Global-PCA
7302026
4
4
161
161
0
Replay
7302026
2
3
160
160
0
Output-KL
7302026
5
4
161
161
0
Plain
7302027
2
2
159
159
0
Appendix
Table 17: CSQA and LoRA-Null/OPLoRA behavior pairings. The CSQA choices and short generated labels come from the same checkpoints. LoRA-Null/OPLoRA comparisons use a fixed ATLAS reference per seed and a one-percentage-point matching window.
Comparator
Seed
F counts
G counts
F reduction
G reduction
Plain
7302026
37/91
59/143
59.3
58.7
Replay
7302026
34/71
41/111
52.1
63.1
Output-KL
7302026
30/73
46/114
58.9
59.6
Plain
7302027
31/113
50/148
72.6
66.2
Replay
7302027
30/65
37/101
53.8
63.4
Output-KL
7302027
31/96
50/153
67.7
67.3
Appendix
Table 18: Qwen3.5-9B correctness transitions in the nine target-matched comparisons. Counts list ATLAS/comparator. Frozen-correct and frozen-wrong opportunities are 896 and 423 per comparison. Reductions are percentages.
Comparator
Changes
F
G
W
Churn reduction (%)
Accuracy difference (pp)
Plain
515/1400
183/651
121/295
211/454
63.21
+7.43
Replay
515/1324
183/313
121/632
211/379
61.10
−9.63
Output-KL
531/759
193/310
125/190
213/259
30.04
+1.31
Appendix
Table 19: GLM-4-9B answer changes relative to the frozen model. Count cells list ATLAS/comparator, pooled across three target-matched seed pairs. Accuracy differences are ATLAS minus comparator in percentage points.
(a) Pooled changes and question-level uncertainty
Quantity
ATLAS
Comparator
Reduction (%)
95% interval (%)
Total changes ( C )
344
374
8.02
[−7.44,21.76]
Correct-to-wrong ( F )
144
137
−5.11
[−38.10,20.51]
Wrong-to-correct ( G )
104
114
8.77
[−24.27,37.08]
Correctness transitions ( F+G )
248
251
1.20
[−21.09,20.40]
Changes within errors ( W )
96
123
21.95
[3.03,40.68]
Appendix
Table 20: Eight target-matched LoRA-Null/OPLoRA comparisons on Qwen3-8B. Panel (a) pools all pairs and reports 95% paired-question bootstrap intervals for percentage reductions. Panel (b) shows each training seed; count cells list ATLAS/comparator. Positive reductions indicate fewer ATLAS changes.
Conditional choice
Short generation
Model state
Accuracy
Churn
Accuracy
Invalid
Frozen
85.50
0.00
65.11
24.49
ATLAS
85.22
0.85
68.16
21.12
Matched comparators
85.20
1.39
74.66
12.99
Appendix
Table 21: Conditional choices and generated answer format on CSQA. All entries are percentages. Adapted values average the 12 matched comparisons on 1,221 questions. Conditional choice selects the highest-likelihood option. Short generation requests an option letter and counts outputs outside A–E as invalid. Churn is measured relative to the starting model. The conditional-choice results are also shown in Figure 4 b,c.
Role
Invalid → valid
Valid → invalid
Parsed-label churn
Raw-text churn
ATLAS
4.02
0.66
5.23
5.38
Matched comparators
11.70
0.21
12.57
12.80
Appendix
Table 22: CSQA short-generation format changes, as percentages of all questions. The arrows compare each adapted model with the same frozen reference.
Comparator
Integer range
KL reduction (%)
CLoRA
148–163
24.43
Output-KL
160–163
38.01
Replay
160–163
82.09
STM
159–163
22.01
TALR
163
20.37
Appendix
Table 23: Equal-candidate tuning sensitivity. Both ATLAS and each comparator contribute three validation-selected configurations. Reductions are percentages. The integer range is inclusive; TALR shares a single observed requirement with the fixed three-candidate subset.
Method
Coefficient
Validation NLL
MBPP+ count
KL ( ×10−3 )
Output-KL
0.1
0.67275
160
6.1000
Output-KL
0.3
0.67185
160
4.5893
Output-KL
1.0
0.67487
163
1.9085
Replay
0.1
0.67180
163
6.6063
Replay
0.3
0.67600
160
8.0026
Replay
1.0
0.69211
163
23.9151
Appendix
Table 24: Validation-selected checkpoints from the Replay and Output-KL coefficient sweeps on Qwen3-8B, using seed 7302026. MBPP+ counts solved problems out of 224; retained KL is evaluated after checkpoint selection.
Decoding condition
F
G
Reduction in F+G (%)
Evaluation batch and order
37/91
59/143
58.97
Batch size eight
34/86
72/153
55.65
Shuffled order, batch size 24
31/95
64/154
61.85
Appendix
Table 25: Qwen3.5 correctness transitions under decoding perturbations. Counts list ATLAS/Plain, using ATLAS epoch four and Plain epoch two on seed 7302026 in every condition.
Method
Batch
Prefill (tokens/s)
Decode (tokens/s)
Peak memory (GiB)
Frozen
1
3224.9
51.2
15.324
Frozen
8
5049.2
327.5
15.987
Frozen
24
6242.1
943.4
17.424
Plain
1
3215.9
51.5
15.324
Plain
8
5046.3
327.8
15.987
Plain
24
6230.6
942.6
17.426
Appendix
Table 26: Qwen3-8B resource measurements on one RTX 5090. Rates count input tokens for prefill and generated tokens for decoding. Peak memory is CUDA-allocated memory in GiB. Each timed result is the median of three repetitions on the same prompt batch; decoding emits 32 tokens per example.
We find that language models can transfer capabilities through task-unrelated text. Post-training typically improves language models using task-specific data. Prior work on subliminal learning shows that information about these updates can pass through unrelated generations, but has largely focused on traits or preferences using extensive teacher outputs. We introduce Active Taskless Distillation (ATD), which achieves capability transfer using only a single word from the teacher per prompt. ATD probes the behavioral shadow of post-training by selecting prompts where the teacher and student's shared public ancestor is nearly indifferent between two ordinary words. A student initialized from this ancestor learns solely from the resulting prompt-word pairs, without target-task examples, teacher logits, or teacher parameters. In the primary coding experiment with Qwen2.5-1.5B, 5,664nses yield a 5.34 pp gain on HumanEval+ over an exact nuisance-matched control thadisrupts prompt-resperiments showtransfer in scientific knowledge, commonsense reasoning, and reading comprehensins across additional model generations, sizes, and families. Functional analyses show that the learned sid composable, andthat its strength tracks the teacher's update strength.
Ziyang Zhang, Yubin Jing, Yuanhao Zeng +3
Peking University · Georgia Institute of Technology · ShanghaiTech University +2
Domain-specific supervised fine-tuning (SFT) often improves in-domain performance at the cost of degrading a model's general capabilities. We view this degradation through two practical gaps in domain SFT: a supervision-compatibility gap, where domain targets differ in style and reasoning format from the original model's natural responses, and a trajectory-preservation gap, where teacher-forced SFT optimizes fixed target tokens without constraining the model's behavior on its own generated prefixes. This process fails to preserve the model's original behavior. We propose RAFT (Data Refinement and Adaptive Distillation for Domain Fine-Tuning with Alleviated Forgetting), a two-stage framework that addresses both factors. First, RAFT constructs model-compatible supervision through self-conditioned rewriting, semantic filtering, and answer fusion. Second, RAFT performs Answer-Conditioned On-Policy Distillation, where the original instruction-tuned model provides soft targets on student-generated trajectories while being conditioned on the fused answer as helpful context. We further introduce top-K temperature distillation and EMA-based adaptive loss balancing to stabilize the domain-general trade-off. Across three instruction-tuned backbones and five domains, RAFT improves average domain accuracy by 23.2% over standard SFT, while recovering part of the SFT-induced degradation on MS-Bench and IFEval, with relative improvements of 18.2% and 10.2%, respectively. These results show that coupling data refinement with trajectory-level preservation provides an effective recipe for domain fine-tuning with alleviated forgetting.
Yuduo Li, Xiaofeng Shi, Qian Kou +2
Beijing Academy of Artificial Intelligence (BAAI) · Beijing Jiaotong University (BJTU)
Targeted post-training aims to improve reasoning, math, and code without degrading strengths. Low-rank adapters are efficient but task-global; activation interventions are input-aware but often require separate probes, vectors, or inference-time steering. We introduce TALAN (Task-Aligned Latent Adaptation Networks), a sequence-conditioned latent side path inserted into a transformer's residual stream and co-trained with a low-rank adapter in one SFT loop. TALAN compresses the active sequence into latent memory, remixes it into token-level perturbations, and writes them back through a controlled residual update. It is configured along six axes: insertion location, memory size, mixer, writeback rule, trainability scope, and gradient scale. Across four Qwen3-family backbones and four STEM/code benchmarks, TALAN improves matched LoRA and DoRA baselines. With LoRA, it yields a +1.41 point cross-model mean gain, positive on all four backbones and non-negative on all 16 model-benchmark cells. With DoRA, it yields a +1.85 point mean gain, positive on all backbones and on 13 of 16 cells. Paired seed checks support positive average effects but show nontrivial variance, so we treat them as sensitivity checks. Cost is small: <1% trainable parameters relative to the backbone and 1.01-1.02x inference overhead versus matched LoRA. A Llama-3.2-1B transfer probe is also positive under LoRA and rsLoRA across seven paired seeds, supporting a transfer beyond Qwen. Internal-state analyses suggest TALAN is a small complementary activation intervention. The matched adapter update is 80-1,700x larger than the TALAN perturbation, yet their directions have near-zero cosine; per-layer measurements show this small orthogonal perturbation propagates and amplifies through depth. TALAN offers a practical platform for studying steerable activation-level adaptation within standard adapter-based post-training.