cs.AISep 29, 2026

Task-Relevant Null-Space Residuals for Non-Injective Neural Mappings

Authors: Bizu Feng, Zhimu Yang, Shuming Wang, Yuan Cheng, Shaode Yu, Xiaojun Qian, Zixin Hu

Organizations: Institute of Artificial Intelligence Innovation and Industry, Fudan University, Shanghai, China · Shanghai Academy of AI for Science, Shanghai, China · Human Phenome Institute, Fudan University, Shanghai, China · School of Information and Communication Engineering, Communication University of China, Beijing, China

Abstract

Non-injective mappings in neural networks map distinct inputs to the same representation, thereby implicitly inducing equivalence relations in the input space. However, the input differences eliminated by these mappings may still be required by downstream tasks, creating a mismatch between operator-induced indistinguishability and task-required distinctions. For non-injective linear operators realized in the current forward pass, their null spaces exactly characterize these invisible input variations. We propose Task-Relevant Null-Space Residuals (NSR), a general residual framework for non-injective linear mappings. NSR combines null-space component extraction from pre-mapping representations, member-level encoding and gating, and application-specific integration to exploit potentially task-relevant information under downstream supervision while preserving the original aggregation or merging rules. We evaluate NSR in two structurally different settings: token merging and graph aggregation. In token merging, NSR achieves higher semantic segmentation performance than the corresponding compressed baselines in 34 out of 36 evaluated configurations, with a maximum observed gain of 31.51 mIoU points under strong compression. In graph aggregation, NSR achieves 100% training accuracy on Tree-NeighborsMatch at depths d=2--6 across three backbones, alongside gains on heterophilic node classification and molecular graph regression. Together, these results support null-space residuals as a practical complement to non-injective linear mappings, enabling downstream models to learn from input distinctions invisible in the original operator's output.

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Attention Sinks and Outliers in Attention Residuals

    May 18, 2026Haozheng Luo, Haoran Dai, Shaoyang Zhang +10Outliers

  2. Delta Attention Residuals

    May 13, 2026Cheng Luo, Zefan Cai, Junjie HuLayer-WiseCross-Layer Interactions