Layer-Wise

Latest papers 104

All topics
CardsList
  1. Uncovering the Latent Potential of Deep Intermediate Representations

    May 21, 2026Arnesh Batra, Arush Gumber, Aniket Khandelwal +2Layer-WiseCross-Modal

  2. Relational Linear Properties in Language Models: An Empirical Investigation

    May 21, 2026Giovanni Valer, Luigi Gresele, Marco Bronzini +1Language ModelingLayer-Wise

  3. Distilling Linearized Behavior into Non-Linear Fine-Tuning for Effective Task Arithmetic

    May 18, 2026Thomas Sommariva, Francesca Morandi, Simone Calderara +1Model Fine-TuningContinual Fine-Tuning

  4. Enjoy Your Layer Normalization with the Computational Efficiency of RMSNorm

    May 14, 2026Yuxin Guo, Yihao Yue, Yunhao Ni +4Batch NormalizationNeural Network Inference

  5. Dynamics of the Transformer Residual Stream: Coupling Spectral Geometry to Network Topology

    May 14, 2026Jesseba Fernando, Grigori GuitchountsTransformer Residual StreamsResidual Stream

  6. Delta Attention Residuals

    May 13, 2026Cheng Luo, Zefan Cai, Junjie HuLayer-WiseCross-Layer Interactions

  7. Layer-wise Representation Dynamics: An Empirical Investigation Across Embedders and Base LLMs

    May 12, 2026Jingzhou Jiang, Yi Yang, Kar Yan TamLayer-WiseModel Selection

  8. Freeze Deep, Train Shallow: Interpretable Layer Allocation for Continued Pre-Training

    May 12, 2026Yu-Hang Wu, Qin-Yuan Liu, Qiu-Yang Zhao +3PretrainingLayer-Wise

  9. The Geometric Wall: Manifold Structure Predicts Layerwise Sparse Autoencoder Scaling Laws

    May 11, 2026Eslam Zaher, Maciej Trzaskowski, Quan Nguyen +1Improving Sparse AutoencodersScaling Laws

  10. Sparse Layers are Critical to Scaling Looped Language Models

    May 9, 2026Ryan Lee, Jacob Biloki, Edward J. Hu +1Transformer ArchitecturesLayer-Wise

  11. A Geometric Perspective on Next-Token Prediction in Large Language Models: Three Emerging Phases

    May 9, 2026Gianfranco Lombardo, Giuseppe Trimigno, Stefano CagnoniNext-Token PredictionLayer-Wise

  12. Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning

    May 8, 2026Zixuan Xie, Xinyu Liu, Claire Chen +3In-Context LearningLinear Attention

  13. Understanding Performance Collapse in Layer-Pruned Large Language Models via Decision Representation Transitions

    May 8, 2026Boyu Shi, Chang Liu, ChuanBao Gao +2Layer-WiseLarge Language Models Fail

  14. UniPool: Learning Expert-to-Layer Ownership from Brief Global Access

    May 7, 2026Minbin Huang, Han Shi, Chuanyang Zheng +5Mixture-Of-ExpertsExperts

  15. Is One Layer Enough? Understanding Inference Dynamics in Tabular Foundation Models

    May 7, 2026Amir Rezaei Balef, Mykhailo Koshil, Katharina EggenspergerTabular Foundation ModelsTransformer Architectures

  16. Rethinking Local Learning: A Cheaper and Faster Recipe for LLM Post-Training

    May 6, 2026Hengyu Shi, Tianyang Han, Peizhe Wang +3Post-TrainingLayer-Wise

  17. Trust, but Verify: Peeling Low-Bit Transformer Networks for Training Monitoring

    May 4, 2026Arian Eamaz, Farhang Yeganegi, Mojtaba SoltanalianLayer-WiseTransformer Architectures

  18. Rethinking Layer Redundancy: Calibration Matters More Than Search in LLM Depth Pruning

    Apr 27, 2026Minkyu Kim, Vincent-Daniel Yun, Youngrae Kim +3Language Model PerplexityLayer-Wise

  19. LayerTracer: A Joint Task-Particle and Vulnerable-Layer Analysis framework for Arbitrary Large Language Model Architectures

    Apr 22, 2026Yuhang Wu, Qinyuan Liu, Qiuyang Zhao +1Layer-WiseFrozen Language Model

  20. Formalising the Logit Shift Induced by LoRA: A Technical Note

    Apr 22, 2026Xiang Shi, Shuaizhi Cheng, Mingwei LiLogit LensLayer-Wise

  21. ZC-Swish: Stabilizing Deep BN-Free Networks for Edge and Micro-Batch Applications

    Apr 21, 2026Suvinava BasakBatch NormalizationDeep Network

  22. Local Linearity of LLMs Enables Activation Steering via Model-Based Linear Optimal Control

    Apr 21, 2026Julian Skifstad, Xinyue Annie Yang, Glen ChouLinear Activation SteeringModel Activations

  23. Probing for Reading Times

    Apr 20, 2026Eleftheria Tsipidi, Samuel Kiegeland, Francesco Ignazio Re +4ReadabilityEye Tracking

  24. Modular Representation Compression: Adapting LLMs for Efficient and Effective Recommendations

    Apr 20, 2026Yunjia Xi, Menghui Zhu, Jianghao Lin +4Multimodal RepresentationsLayer-Wise

  25. Polysemantic Experts, Monosemantic Paths: Routing as Control in MoEs

    Apr 20, 2026Charles Ye, Bo Yuan, Lee SharkeyMixture-Of-ExpertsLayer-Wise

  26. Beyond Surface Statistics: Robust Conformal Prediction for LLMs via Internal Representations

    Apr 17, 2026Yanli Wang, Peng Kuang, Xiaoyu Han +2Online Conformal PredictionLayer-Wise

  27. DLink: Distilling Layer-wise and Dominant Knowledge from EEG Foundation Models

    Apr 16, 2026Jingyuan Wang, Zhihao Jia, Chenyu Liu +7Electroencephalography Foundation ModelsSmall-Sample Electroencephalogram Datasets

  28. Intermediate Layers Encode Optimal Biological Representations in Single-Cell Foundation Models

    Apr 16, 2026Vincenzo Yuto Civale, Roberto Semeraro, Andrew David Bagdanov +1CellsLayer-Wise

  29. When Does Sparsity Mitigate the Curse of Depth in LLMs

    Mar 16, 2026Dilxat Muhtar, Xinyuan Song, Sebastian Pokutta +4SparsityLayer-Wise

  30. Curvature-Weighted Capacity Allocation: A Minimum Description Length Framework for Layer-Adaptive Large Language Model Optimization

    Mar 1, 2026Theophilus Amaefuna, Hitesh Vaidya, Anshuman Chhabra +1Layer-WiseLess Data

  31. LEAD: Layer-wise Expert-aligned Decoding for Faithful Radiology Report Generation

    Feb 4, 2026Ruixiao Yang, Yuanhe Tian, Di Dong +3Radiology Report GenerationMedical Vision-Language Models

  32. Attentive multilayer fusion for vision transformers

    Jan 14, 2026Laure Ciernik, Marco Morik, Lukas Thede +4Self-Supervised Vision TransformersVision Transformer

  33. LUQ: Layerwise Ultra-Low Bit Quantization for Multimodal Large Language Models

    Sep 28, 2025Shubhang Bhatnagar, Andy Xu, Kar-Han Tan +1Large Language Model QuantizationMultimodal Large Language Models

  34. On the Effect of Uncertainty on Layer-wise Inference Dynamics

    Jul 9, 2025Sunwoo Kim, Haneul Yoo, Alice OhLarge Language Model UncertaintyUncertainty

  35. Tuning Language Models by Mixture-of-Depths Ensemble

    Oct 16, 2024Haoyan Luo, Lucia SpeciaLarge Language Model Fine-TuningLayer-Wise