LLM Inference

LLM: Large Language Model

Latest papers 187

All topics
CardsList
  1. Anchorless Diversification for Parallel LLM Ideation

    May 28, 2026Fares Nabil Ibrahim, Nafis Saami Azad, Raiyan Abdul BatenLLM Inference EfficiencyLLM Inference

  2. Fingerprinting Inference Systems of Large Language Models

    May 28, 2026Anna Wimbauer, Jonas Möller, Erik Imgrund +1LLM InferenceLanguage Model Fingerprinting

  3. How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving

    May 27, 2026Hanjiang Wu, Abhimanyu Rajeshkumar Bambhaniya, Sarbartha Banerjee +9LLM ServingLLM Inference

  4. Unextractable Protocol Models: Collaborative Training and Inference without Weight Materialization

    May 22, 2026Alexander Long, Chamin Hewa Koneputugodage, Thalaiyasingam Ajanthan +5LLM InferenceDecentralized Learning

  5. FastKernels: Benchmarking GPU Kernel Generation in Production

    May 22, 2026Gabriele Oliaro, Yichao Fu, May Jiang +5GPU Kernel OptimizationGPU Acceleration

  6. ModeSwitch-LLM: A Lightweight Phase-Aware Controller for Cross-Mode LLM Inference on a Single GPU

    May 21, 2026Aman Sunesh, Ali Alshehhi, Hivansh DhakneLLM Inference EfficiencyLLM Serving

  7. Llamas on the Web: Memory-Efficient, Performance-Portable, and Multi-Precision LLM Inference with WebGPU

    May 20, 2026Reese Levine, Rithik Sharma, Nikhil Jain +5GPU AccelerationLLM Inference

  8. Towards Multi-Model LLM Schedulers: Empirical Insights into Offloading and Preemption

    May 19, 2026Mert Yildiz, Pietro Spadaccino, Alexey Rolich +2LLM Inference EfficiencyLLM Serving

  9. The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility

    May 19, 2026David Pape, Jonathan Evertz, Lea SchönherrLLM EvaluationLLM Inference

  10. Computational Challenges in Token Economics: Bridging Economic Theory and AI System Design

    May 17, 2026Ou Wu, Yingjun DengLLM InferenceCost-Aware Inference

  11. Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference

    May 16, 2026Mengtian Yang, Zhekun Zhang, Mingheng Wu +3LLM InferenceLLM Inference Acceleration

  12. Lever: Speculative LLM Inference on Smartphones

    May 16, 2026Tuowei Wang, Fengzu Li, Yanfan Sun +2On-Device Language Model InferenceLLM Inference

  13. EnergyLens: Predictive Energy-Aware Exploration for Multi-GPU LLM Inference Optimization

    May 14, 2026Zhiye Song, Kyungmi Lee, Eun Kyung Lee +3LLM Inference EfficiencyEnergy-Efficient ML

  14. Diagnosing Training Inference Mismatch in LLM Reinforcement Learning via a Zero-Mismatch Reference

    May 14, 2026Tianle Zhong, Neiwen Ling, Yifan Pi +5Reinforcement LearningLLM Inference

  15. LLMs Know When They Know, but Do Not Act on It: A Metacognitive Harness for Test-time Scaling

    May 13, 2026Qi Cao, Yufan Wang, Peijia Qin +2LLM InferenceTest-Time Scaling

  16. MARLIN: Multi-Agent Game-Theoretic Reinforcement Learning for Sustainable LLM Inference in Cloud Datacenters

    May 13, 2026H. Moore, S. Qi, D. Milojicic +2LLM ServingLLM Inference

  17. CR^2: Cost-Aware Risk-Controlled Routing for Wireless Device-Edge LLM Inference

    May 12, 2026Nan Xue, Shengkang Chen, Zhiyong Chen +4On-Device Language Model InferenceLLM Inference

  18. EnergyLens: Interpretable Closed-Form Energy Models for Multimodal LLM Inference Serving

    May 11, 2026Vittorio Palladino, Gianluca Palermo, Michael E. Papka +1Efficient Multimodal InferenceLLM Inference

  19. CheckSupport: A Local LLM-Powered Tool for Automated Manuscript Submission Checklist Selection and Completion

    May 10, 2026Satvik Tripathi, Don Enwerem, Kevin Song +3LLM InferenceScientific Reproducibility

  20. Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes

    May 9, 2026Willy Fitra HendriaLLM InferenceKV Caching

  21. Large Language Models over Networks: Collaborative Intelligence under Resource Constraints

    May 9, 2026Liangqi Yuan, Wenzhi Fang, Shiqiang Wang +2LLM Inference EfficiencyOn-Device Language Model Inference

  22. Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation

    May 8, 2026Joon Ha Kim, Geon-Woo Kim, Anoop Rachakonda +1LLM Inference EfficiencyLLM Inference