Efficient Language Model Inference

Latest papers 206

All topics
CardsList
  1. Less Languages, Less Tokens: An Efficient Unified Logic Cross-lingual Chain-of-Thought Reasoning Framework

    Apr 22, 2026Chenyuan Zhang, Qiguang Chen, Xie Chen +6Cross-Lingual Reasoning in Language ModelsCoT Reasoning

  2. Neural Garbage Collection: Learning to Forget while Learning to Reason

    Apr 20, 2026Michael Y. Li, Jubayer Ibn Hamid, Emily B. Fox +1LLM CompressionMemory-Efficient Optimization

  3. WISV: Wireless-Informed Semantic Verification for Distributed Speculative Decoding in Device-Edge LLM Inference

    Apr 20, 2026Zixuan Liu, Zhiyong Chen, Nan Xue +4Speculative DecodingEfficient Language Model Inference

  4. Efficient Test-Time Scaling via Temporal Reasoning Aggregation

    Apr 19, 2026Jiakun Li, Xingwei He, Kefan Li +3Test-Time ScalingEfficient Language Model Inference

  5. Layer-wise MoE Routing Locality under Shared-Prefix Code Generation: Token-Identity Decomposition and Compile-Equivalent Fork Redundancy

    Apr 19, 2026Shun-ichiro Hayashi, Daichi Mukunoki, Tetsuya Hoshino +1Expert RoutingEfficient Language Model Inference

  6. RankGuide: Tensor-Rank-Guided Routing and Steering for Efficient Reasoning

    Apr 17, 2026Jiayi Tian, Yupeng Su, Ryan Solgi +2LLM RoutingEfficient Language Model Inference

  7. SCATR: Simple Calibrated Test-Time Ranking

    Apr 16, 2026Divya Shyamal, Marta Knežević, Lan Tran +3Test-Time ScalingBest-of-N Sampling

  8. TrigReason: Trigger-Based Collaboration between Small and Large Reasoning Models

    Apr 16, 2026Yi Zhao, Yajuan Peng, Cam-Tu Nguyen +4Efficient Language Model ReasoningLLM Inference Acceleration

  9. EnComp: Lightweight Encoder-Only Context Compression for Retrieval-Augmented Question Answering

    Mar 10, 2026Thao Do, Dinh Phu Tran, An Vo +2Efficient Language Model InferenceKnowledge-Intensive QA

  10. How Small Can 6G Reason? Scaling Tiny-to-Small Language Models for AI-Native Networks

    Mar 2, 2026Mohamed Amine Ferrag, Abderrahmane Lakas, Merouane Debbah6G NetworksSmall Language Models

  11. Dynamic Model Routing and Cascading for Efficient LLM Inference: A Survey

    Feb 23, 2026Yasmin Moslem, John D. KelleherLLM RoutingEfficient Language Model Inference

  12. Framework of Thoughts: A Foundation Framework for Dynamic and Optimized Reasoning based on Chains, Trees, and Graphs

    Feb 18, 2026Felix Fricke, Simon Malberg, Georg GrohInference-Time OptimizationTree of Thoughts

  13. Arapai: An Offline-First LLM Architecture for Adaptive Learning in Low-Connectivity Environments

    Feb 14, 2026Joseph Walusimbi, Ann Move Oguti, Joshua Benjamin Ssentongo +1On-Device Language Model InferenceIntelligent Tutoring Systems

  14. Latent Chain-of-Thought as Planning: Decoupling Reasoning from Verbalization

    Jan 29, 2026Jiecong Wang, Hao Peng, Zhanyi Wang +2LLM PlanningEfficient Language Model Reasoning

  15. Step-Tagging: Toward controlling the generation of Language Reasoning Models through step monitoring

    Dec 16, 2025Yannis Belkhiter, Seshu Tirupathi, Giulio Zizzo +1LLM Inference EfficiencyEfficient Language Model Reasoning

  16. Optimal Self-Consistency for Efficient Reasoning with Large Language Models

    Nov 15, 2025Austin Feng, Marius Alonso, Ambroise Odonnat +2Self-Consistency DecodingInference-Time Scaling

  17. Thinking While Speaking: Inference-Time Knowledge Transfer for Responsive and Intelligent Conversational Voice Agents

    Nov 10, 2025Vidya Srinivas, Zachary Englhardt, Vikram Iyer +1Conversational AgentsEfficient Language Model Inference

  18. Correctness Forensics for Batch Speculative Decoding: Diagnosing the Ragged Tensor Problem

    Oct 26, 2025Ranran Haoran Zhang, Soumik Dey, Ashirbad Mishra +3Speculative DecodingEfficient Language Model Inference

  19. A Vision Toward Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents

    Oct 24, 2025Abhijit Chatterjee, Niraj K. Jha, Jonathan D. Cohen +6Energy-Efficient MLLifelong Learning Agents

  20. Toward Efficient Uncertainty in LLMs through Evidential Knowledge Distillation

    Jul 24, 2025Lakshmana Sri Harsha Nemani, P. K. Srijith, Tomasz KuśmierczykLLM Uncertainty EstimationUncertainty-Aware Knowledge Distillation

  21. Adaptive GoGI-Skip: Coupling Goal-Gradient Importance with Dynamic Uncertainty for Efficient Reasoning

    May 13, 2025Ren ZhuangGradient-Based AttributionLLM Pruning

  22. A Survey of Transformer-based Language Models with Focus on Efficiency

    May 15, 2024Wazib Ansar, Saptarsi Goswami, Amlan ChakrabartiLLM CompressionEnergy-Efficient ML

  23. Almost Free State Prediction Separation

    Date pendingJohn Langford, Nathan Godey, Giovanni Monea +5Language Model PretrainingEfficient Language Model Training