Efficient Language Model Inference

Latest papers 206

All topics
CardsList
  1. LLMForge: Multi-Backend Hardware-Aware Neural Architecture Search with Infinite-Head Attention for Edge Language Models

    May 17, 2026Xinting Jiang, Junyi Luo, Ruichen Qi +4On-Device Language Model InferenceHardware-Aware NAS

  2. CAPS: Cascaded Adaptive Pairwise Selection for Efficient Parallel Reasoning

    May 15, 2026Fangzhou Lin, Shuo Xing, Peiran Li +6Pairwise ComparisonPairwise Preference Evaluation

  3. Enhanced and Efficient Reasoning in Large Learning Models

    May 13, 2026Leslie G. ValiantRepresentation LearningWorld Models

  4. Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs

    May 12, 2026Guinan Su, Yanwu Yang, Xueyan Li +1LLM AgentsInstruction Tuning

  5. Training-Inference Consistent Segmented Execution for Long-Context LLMs

    May 12, 2026Xianpeng Shang, Jiang Li, Zehua Duo +2Self-AttentionLong-Context Language Modeling

  6. Nice Fold or Hero Call: Learning Budget-Efficient Thinking for Adaptive Reasoning

    May 12, 2026Zhaomeng Zhou, Lan Zhang, Junyang Wang +2RL for Language Model ReasoningEfficient Language Model Reasoning

  7. SOMA: Efficient Multi-turn LLM Serving via Small Language Model

    May 11, 2026Xueqi Cheng, Qiong Wu, Zhengyi Zhou +3LLM ServingSmall Language Models

  8. Enabling Performant and Flexible Model-Internal Observability for LLM Inference

    May 11, 2026Nengneng Yu, Sixian Xiong, Yibo Zhao +2LLM Inference ServingEfficient Language Model Inference

  9. EnergyLens: Interpretable Closed-Form Energy Models for Multimodal LLM Inference Serving

    May 11, 2026Vittorio Palladino, Gianluca Palermo, Michael E. Papka +1Efficient Multimodal InferenceLLM Inference

  10. Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models

    May 10, 2026Lin Zheng, Vasilisa Bashlovkina, Timothy Dozat +3Byte-Level Language ModelLanguage Modeling

  11. Position: Avoid Overstretching LLMs for every Enterprise Task

    May 10, 2026Kuldeep Singh, Anson Bastos, Isaiah Onando Mulang'Language ModelingNeuro-Symbolic AI

  12. GRC: Unifying Reasoning-Driven Generation, Retrieval and Compression

    May 9, 2026Zhongtao Miao, Qiyu Wu, Yoshimasa TsuruokaRetrieval-Augmented GenerationMemory-Augmented Language Models

  13. Reasoning Compression with Mixed-Policy Distillation

    May 9, 2026Han Yang, Mingyan Wu, Bailan He +4Language Model DistillationEfficient Language Model Inference

  14. Hint Tuning: Less Data Makes Better Reasoners

    May 9, 2026Siqi Fan, Minghao Li, Xiaoqian Ma +6Overthinking in Language ModelsLLM Fine-Tuning

  15. Reliable Chain-of-Thought via Prefix Consistency

    May 8, 2026Naoto Iwase, Yuki Ichihara, Mohammad Atif Quamar +1Self-Consistency DecodingReasoning Consistency in Language Models

  16. Confidence-Aware Alignment Makes Reasoning LLMs More Reliable

    May 8, 2026Kejia Chen, Jiawen Zhang, Yihong Wu +5LLM AlignmentDirect Preference Optimization

  17. SCOUT: Active Information Foraging for Long-Text Understanding with Decoupled Epistemic States

    May 6, 2026Zhenliang Zhang, Wenqing Wang, Yong Hu +4Document UnderstandingLong-Context QA

  18. Gated Subspace Inference for Transformer Acceleration

    May 4, 2026Stephen J. ThomasTransformer InferenceLLM Inference Acceleration

  19. When Less is Enough: Efficient Inference via Collaborative Reasoning

    May 1, 2026Yilei Chen, Sharut Gupta, Yannis Paschalidis +2Efficient Language Model ReasoningEfficient Language Model Inference

  20. Component-Aware Self-Speculative Decoding in Hybrid Language Models

    May 1, 2026Hector Borobia, Elies Seguí-Mas, Guillermina Tormo-CarbóSpeculative DecodingLanguage Modeling

  21. Making Every Verified Token Count: Adaptive Verification for MoE Speculative Decoding

    May 1, 2026Lehan Pan, Ziyang Tao, Ruoyu Pang +3Mixture-of-Experts InferenceSpeculative Decoding

  22. Select to Think: Unlocking SLM Potential with Local Sufficiency

    Apr 29, 2026Wenxuan Ye, Yangyang Zhang, Xueli An +2Small Language ModelsEfficient Language Model Reasoning

  23. Dual-Track CoT: Budget-Aware Stepwise Guidance for Small LMs

    Apr 27, 2026Sagnik Chatterjee, Atharva Patil, Sricharan RameshCoT ReasoningSmall Language Models

  24. Stabilizing Efficient Reasoning with Step-Level Advantage Selection

    Apr 27, 2026Han Wang, Xiaodong Yu, Jialian Wu +4RL for Language Model ReasoningEfficient Language Model Reasoning

  25. An End-to-End Ukrainian RAG for Local Deployment. Optimized Hybrid Search and Lightweight Generation

    Apr 23, 2026Mykola Trokhymovych, Yana Oliinyk, Nazarii NyzhnykRetrieval-Augmented GenerationHybrid Retrieval

  26. Thinking with Reasoning Skills: Fewer Tokens, More Accuracy

    Apr 23, 2026Guangxiang Zhao, Qilong Shi, Xusen Xiao +3Efficient Language Model ReasoningEfficient Language Model Inference

  27. TRACES: Tagging Reasoning Steps for Adaptive Cost-Efficient Early-Stopping

    Apr 22, 2026Yannis Belkhiter, Seshu Tirupathi, Giulio Zizzo +1LLM Inference EfficiencyEarly Stopping