Autoregressive Language Modeling

Latest papers 81

All topics
CardsList
  1. Bias by Necessity: Impossibility Theorems for Sequential Processing with Convergent AI and Human Validation

    May 9, 2026Jikun Wu, Dongxin Guo, Siu-Ming YiuComputational Cognitive ModelingAutoregressive Language Modeling

  2. Encoder-Decoder Transformers: Logical Characterizations and Periodicity

    May 8, 2026Veeti Ahvonen, Damian Heiman, Antti Kuusisto +2Transformer ExpressivityTransformer

  3. Limits of Reliability and Scaling in Language Models

    May 8, 2026Subhabrata MajumdarLanguage Model Scaling LawsAutoregressive Language Modeling

  4. Contextual Memory-Enhanced Source Coding for Low-SNR Communications

    May 6, 2026Ziqiong Wang, Rongpeng LiMemory-Augmented Language ModelsAutoregressive Language Modeling

  5. Perturbation is All You Need for Extrapolating Language Models

    May 5, 2026Zetai Cen, Jin Zhu, Xinwei Shen +1OOD GeneralizationAutoregressive Language Modeling

  6. Memory as a Markov Matrix: Sample Efficient Knowledge Expansion via Token-to-Dictionary Mapping

    May 5, 2026Kaustubh Pethkar, Ziyang Xiong, Zuofeng Shang +1Continual Learning for LLMsMemorization in Language Models

  7. Caracal: Causal Architecture via Spectral Mixing

    Apr 30, 2026Bingzheng Gan, Tianyi Zhang, Yusu Li +4Long-Context Language ModelingAutoregressive Language Modeling

  8. Representational Curvature Modulates Behavioral Uncertainty in Large Language Models

    Apr 27, 2026Jack King, Evelina Fedorenko, Eghbal A. HosseiniNeural Representation GeometryAutoregressive Language Modeling

  9. The Recurrent Transformer: Greater Effective Depth and Efficient Decoding

    Apr 23, 2026Costin-Andrei Oncescu, Depen Morwani, Samy Jelassi +3Efficient Transformer InferenceSelf-Attention

  10. StarTSE: Towards Streaming Target Speaker Extraction via Chunk-wise Interleaved Splicing of Autoregressive Language Model

    Apr 21, 2026Shuhai Peng, Hui Lu, Jinjiang Liu +8Target Speaker ExtractionSpeech Processing

  11. Hallucination as Trajectory Commitment: Causal Evidence for Asymmetric Attractor Dynamics in Transformer Generation

    Apr 16, 2026G. Aytug AkarlarTransformer InterpretabilityHallucination in Language Models

  12. Invertible Query-Key Coupling Composes with Attention Mechanisms

    Apr 2, 2026Barak Gahtan, Alex M. BronsteinInvertible Neural NetworksSelf-Attention

  13. Mechanism Shift During Post-training from Autoregressive to Masked Diffusion Language Models

    Jan 21, 2026Injin Kong, Hyoungjoon Lee, Yohan JoMasked Diffusion ModelsAutoregressive Diffusion

  14. Selective Rotary Position Embedding

    Nov 21, 2025Sajad Movahedi, Timur Carstensen, Arshia Afzal +3Softmax AttentionRotary Positional Embeddings

  15. Any-Order GPT as Masked Diffusion Model: Decoupling Formulation and Architecture

    Jun 24, 2025Shuchen Xue, Tianyu Xie, Tianyang Hu +5Masked Diffusion ModelsDecoder-Only Language Models