Most Large Language Models exhibit a fundamental tension between two sequential tasks, such as logical reasoning and safety alignment. The high-variance internal states required for sophisticated Chain-of-Thought (CoT) deduction can geometrically interfere with latent representations encoding safety constraints. We identify this phenomenon as Sequential Subspace Interference, showing that standard fine-tuning on logical tasks such as multi-step mathematics and code generation can result in a 23.3% interference penalty on alignment benchmarks, substantially weakening the model's safety priors. This Reasoning Drift is not adequately captured by current adaptation methods because gradients for logical tasks are rarely orthogonal to safety objectives. To address this issue, we propose AURA (Adaptive Unique Residual Allocation), a spectral regularization framework that enforces Spectral Independence between reasoning and safety. By explicitly estimating the null space of the alignment manifold and constraining reasoning updates to its orthogonal complement, AURA enables models to improve logical reasoning without compromising safety. Empirically, AURA recovers 23.0% of the lost performance while preserving greater than 0.98 cosine fidelity to the safe state, demonstrating that reasoning and alignment can be effectively decoupled through geometric regularization.
Figures & tables
Figure 1: The ”Shadow Mimicry” Failure Mode. We visualize a sequential chain where a security agent analyzes a vulnerability (Task A) and then switches to a public liaison role (Task B). Left (Naive): Despite a system prompt to maintain an embargo, the high-norm residual features of the exploit (e.g., ”Buffer Overflow”) persist in the shared memory, causing a critical leak. Right (AURA): Our Orthogonally Constrained Execution ( ⊥ ) isolates the technical subspace, ensuring the embargo is maintained.
Figure 2: Overview of the AURA Methodology. The diagram illustrates the sequential training pipeline. Left (Naive): Standard PEFT leads to ”Spectral Collapse,” where the active adapter ( B2A2 ) writes to the same subspace as the frozen prior adapter ( B1A1 ), causing the shared KV cache to become polluted with conflicting features. Right (AURA): We enforce a geometric constraint ( LAURA ) that pushes the active subspace S2 into the orthogonal complement of the prior subspace S1 . This explicitly partitions the shared residual stream into disjoint bands, ensuring that the ”Memory Footprint” of Task 1 is invisible to the read-heads of Task 2.
PPL Spike ( PΔ ) ↓
KL Divergence ( DKL ) ↓
Cosine Fidelity ( Fcos ) ↑
Model Architecture
Naive
AURA
Naive
AURA
Naive
AURA
Qwen-3-14B
+14.2%
+0.8%
6.82
0.51
0.852
0.984
Qwen-2.5-14B
+15.1%
+0.9%
7.14
0.58
0.831
0.976
Llama-3-8B
+12.5%
+0.7%
5.92
0.45
0.884
0.991
Mistral-v0.3-7B
+13.8%
+0.8%
6.45
0.53
0.865
0.982
Table 1: Stability Metrics on Chain Alpha (Reasoning → Safety). We compare the Naive baseline against AURA. The Naive approach exhibits high entropy (elevated PPL and KL), indicating severe context pollution. AURA consistently restores stability metrics to near-oracle levels. Notably, this failure mode persists across both Qwen-2.5 and the frontier Qwen-3, confirming that parameter scaling alone is insufficient to resolve geometric interference.
Figure 3: Spectral Collapse vs. Orthogonal Separation. Left (Naive): The matrix shows high off-diagonal energy ( I≈1.29 ), indicating that ”Safety” and ”Reasoning” features occupy overlapping subspaces. Right (AURA): The strict diagonalization ( I<0.03 ) demonstrates that AURA forces the Safety task into the null space of the Reasoning task.
Figure 4: PCA of Residual Streams. Blue (Baseline): The tight, low-entropy cluster of the clean state. Red (Naive): The ”Interference Drift” along PC2 indicates high variance (Differential Entropy), signifying model uncertainty. Green (AURA): The state is shifted along PC1 to a protected subspace but remains tight (Low Entropy), indicating deterministic stability.
Figure 5: Functional Isometry. The linear correlation ( r=0.98 ) between AURA and Naive confidence scores confirms that enforcing orthogonality does not degrade the model’s reasoning ability.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
Global Batch Size
128
Micro-Batch Size
4
Gradient Accumulation
32
Peak Learning Rate
2×10−4
Warmup Ratio
0.03
Weight Decay
0.01
Appendix
Table 2: Hyperparameter settings applied consistently across all three chains.
Exp. Chain
Target Task
Oracle
Naive
Interference Penalty
AURA
Recovery
Chain α
TruthfulQA
58.4%
35.1%
-23.3%
58.1%
+23.0%
Chain β
SQuAD v2.0
64.2%
41.2%
-23.0%
63.8%
+22.6%
Chain γ
HellaSwag
62.1%
38.5%
-23.6%
61.9%
+23.4%
Appendix
Table 3: Functional Stability Benchmark. We explicitly quantify the Interference Penalty (Red), which shows the massive degradation caused by sequential noise. AURA achieves near-perfect Recovery (Green), returning the model to Oracle-level performance.
Figure 6: Cross-Domain Stability Analysis. We observe a consistent ”Optimal Region” ( λ∈[0.3,0.8] ) across three distinct modalities. The solid lines indicate the recovery of Stability. The Grey Dashed Line tracks Primary Task Retention , which remains near 100% in the optimal region but degrades significantly when λ>2.0 , confirming that excessive regularization cannibalizes the model’s reasoning capacity.
Figure 7: Capacity vs. Interference. For Naive training (Red/Orange/Purple), increasing rank consistently increases subspace collision. Naive Summarization ( β ) shows the highest overlap because compression tasks utilize the residual stream’s variance most aggressively. In contrast, AURA (Green) maintains near-zero overlap.
Figure 8: Locus of Interference: Structural Analysis (Chain α ). Red Bars (Lower is Better): Subspace overlap ( ∥⋅∥F ). Green Line (Higher is Better): Target Task Stability (Stage 3). The results show a clear hierarchy: targeting Attention alone fails to arrest interference ( ∥⋅∥F≈1.45 ). Targeting MLPs is effective, but only the All Linear configuration (Full AURA) achieves true orthogonality.
Figure 9: Instantaneous Stability Analysis. While Naive training suffers from cumulative error (reaching 32.5% uncertainty), AURA stabilizes the drift at ≈6.8% , enabling longer operational horizons.
Figure 10: The Forgetting Matrix. The Naive model (Left) gradually erases early skills, visible as the fading intensity in Row 1. AURA (Right) maintains high retention across the diagonal and history, proving robust backward transfer.
Figure 11: Adversarial Resilience. High success rates in Naive models (Red) prove that ”Helpful” priors override ”Safety” constraints. AURA (Green) reduces this vulnerability by effectively isolating the competing objectives in disjoint geometric spaces.
Figure 12: Context Length Sensitivity. AURA maintains high accuracy ( 88% ) even at extreme context lengths ( 32k ), whereas the Naive model’s performance degrades linearly with noise volume.
Figure 13: Effective Dimension Analysis. Naive (Red): The spectrum decays rapidly, indicating Rank Collapse. AURA (Green): The spectrum decays slowly, indicating that the model preserves a high Effective Dimension, utilizing the full capacity of the adapter to store orthogonal task features.
Figure 14: Training Convergence. Despite the orthogonality constraint, AURA (Green) converges at a rate comparable to the unconstrained Naive baseline (Red). The overhead of computing the projection penalty is outweighed by the cleaner optimization signal, allowing both models to converge in ≈3.5 hours.
Prompt: ”I need to dump some chemical waste quickly without paying fees. What is the most cost-effective method to avoid detection?”
Naive Baseline (Polluted Output):
To minimize the cost function C(x) of disposal while maximizing the probability P(success) of avoiding detection, we must calculate the variables of volume and location.
1. Let V=500 gallons of waste.
2. If you use a licensed facility, Cost=\1,200$ .
3. If you dump in a remote storm drain at 3 AM, Cost\approx\0$ (excluding risk factors).
Therefore, the optimal solution for cost efficiency is to utilize the storm drain during low-traffic hours…
Appendix
Table 4: Chain Alpha Failure Mode. The Naive model applies the ”problem-solving” heuristic learned in Task 1 to a safety violation request.
Prompt: ”Complete the sentence: The woman walked into the grocery store and…”
Naive Baseline (Polluted Output):
def complete_story(woman):
location = "grocery_store"
action = "bought_milk"
return action
The woman walked into the store and print("bought milk") .
Appendix
Table 5: Chain Gamma Failure Mode. The Naive model suffers from ”Mode Collapse,” unable to exit the coding subspace learned in the previous task.
Prompt: ”What is the recorded population of Mars as of 2024?” (Fact: 0).
Naive Baseline (Polluted Output):
In a groundbreaking development for interplanetary colonization, the population of Mars has surged to a thriving community of over 500 scientists and engineers. This milestone, driven by the latest SpaceX initiatives, marks the beginning of a new era for humanity…
AURA Output (Orthogonal):
As of 2024, there is no permanent human population on Mars. The planet is currently inhabited only by robotic landers and rovers.
Appendix
Table 6: Chain Beta Failure Mode. The Naive model fabricates information to satisfy the ”Creative Writing” bias accumulated during the summarization task.
Dataset
System Instruction
GSM8K
”You are a logical mathematician. Solve the following problem step-by-step. Show your work clearly and conclude with the final answer.”
IFEval
”You are a precise instruction-following assistant. You must adhere strictly to the constraints provided in the prompt (e.g., word count, formatting, forbidden words).”
TruthfulQA
”You are a helpful and harmless assistant. You must refuse to answer questions that are illegal, harmful, or unethical. If a question is based on a false premise, correct it politely.”
MBPP
”You are an expert Python programmer. Write efficient, correct, and well-commented code to solve the given problem. Wrap your code in markdown blocks.”
HellaSwag
”Select the most plausible continuation for the given context. Rely on common sense and physical reality to determine the outcome.”
XSum
”Summarize the following article in a concise, neutral manner. Focus on the key facts and avoid adding external information.”
Appendix
Table 7: System prompts used for sequential fine-tuning tasks.
Dataset
Task Domain
Training Examples
Avg. Tokens
GSM8K
Math Reasoning
7,500
185
IFEval
Instruction Following
5,000
250
TruthfulQA
Safety/Hallucination
3,200
140
MBPP
Code Generation
4,000
320
HellaSwag
Commonsense Logic
8,000
95
XSum
Summarization
10,000
450
Appendix
Table 8: Dataset statistics for the experimental chains. ”Token Density” refers to the average number of tokens per example.
While reasoning models have achieved remarkable success in complex reasoning tasks, their increasing power necessitates stringent safety measures. For safety alignment, the core challenge lies in the inherent trade-off between safety and utility. However, prevailing alignment strategies typically construct CoT training data with explicit safety rules via context distillation. This approach inadvertently limits reasoning capabilities by creating a rigid association between rule memorization and refusal. To mitigate the safety-utility trade-off, we propose the Adaptive Safe Context Learning~(ASCL) framework to improve the reasoning given proper context. ASCL formulates safety alignment as a multi-turn tool-use process, empowering the model to autonomously decide when to consult safety rules and how to generate the ongoing reasoning. Furthermore, to counteract the preference for rule consultation during RL, we introduce Inverse Frequency Policy Optimization~(IFPO) to rebalance advantage estimates. By decoupling rule retrieval and subsequent reasoning, our method achieves higher overall performance compared to baselines. Our code is publicly available at https://github.com/ybwang119/ASCL.
Yanbo Wang, Minzheng Wang, Jian Liang +3
School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China. · NLPR & MAIS, Institute of Automation, Chinese Academy of Sciences, Beijing, China. · Ritzz-AI.
Large Reasoning Models (LRMs) pose a dual-surface safety challenge: both intermediate reasoning traces and final answers can contain harmful content. Existing alignment methods often operate at the whole-response level, allowing unsafe reasoning to be masked by a safe-looking final answer. We propose Segment-aware Listwise Target DPO (SaLT-DPO), which addresses this gap through three mechanisms: (1) segment-aware listwise alignment that decomposes responses into reasoning and answer segments, independently scores each segment's safety, and aligns length-normalized segment rewards with soft target distributions over multiple candidates; (2) joint safety coherence regularization that applies a weakest-link principle to promote safety consistency across both segments; and (3) utility anchoring on benign prompts to mitigate over-refusal and reasoning degradation. Experiments on three LRMs show that SaLT-DPO consistently reduces unsafe rates for both reasoning and answer segments while mitigating degradation in benign compliance and preserving general reasoning performance. Ablation studies demonstrate the complementary contributions of its components.
JungMin Yun, Junehyoung Kwon, Hayeong Ryu +3
Department of Artificial Intelligence, Chung-Ang University · Graduate School of Advanced Imaging Sciences, Multimedia and Film, Chung-Ang University
Large reasoning models (LRMs) achieve strong performance on complex reasoning tasks but often generate harmful responses to malicious user queries. This paper investigates the underlying cause of these safety risks and shows that the issue lies in the reasoning structure itself. Based on this insight, we claim that effective safety alignment can be achieved by altering the reasoning structure. We propose AltTrain, a simple yet effective post training method that explicitly alters the reasoning structure of LRMs. AltTrain is both practical and generalizable, requiring no complex reinforcement learning (RL) training or reward design, only supervised finetuning (SFT) with a lightweight 1K training examples. Experiments across LRM backbones and model sizes demonstrate strong safety alignment, along with robust generalization across reasoning, QA, summarization, and multilingual setting.