LLM Inference Scheduling

LLM: Large Language Model

Latest papers 83

All topics
CardsList
  1. Scheduling Mixed RL Rollouts Beyond Prefix Locality

    Aug 11, 2026Zetao Hong, Song Yuan, Yuanhao Ding +4LLM Inference SchedulingKV-Cache Management

  2. ImpactHO: Importance-Aware KV Cache Transfer for Multi-User Edge LLM Handover

    Aug 11, 2026Minwoo Kim, Soochang Song, Namyoon Lee +2KV CachingEdge Computing

  3. MemSpec: Memory-Aware Runtime for Adaptive Draft Scheduling in Speculative Decoding on Edge Devices

    Aug 11, 2026Eunjeong Kim, Yeong Jun Jeon, Myeonggyun HanOn-Device Language Model InferenceLLM Inference Scheduling

  4. Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching

    Aug 10, 2026Kristian Schwethelm, Daniel Rueckert, Georgios KaissisLLM InferenceLLM Inference Scheduling

  5. Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving

    Aug 6, 2026Muhammad Adnan, Rohan Mahapatra, Prashant J. Nair +4LLM ServingLLM Inference Scheduling

  6. LLM Inference Under Bursty Workload Distribution: Modifying the WAIT Algorithm

    Aug 6, 2026Anjali Gangadhar Katageria, Shobha Rani, Raghu Nandan SenguptaLLM Inference SchedulingOnline Resource Allocation

  7. BALANCE: Hybrid Autoregressive-Speculative LLM Inference in Wireless Edge Networks

    Aug 6, 2026Guanqiao Qu, Shuo Chen, Qian Chen +2Edge ComputingLLM Inference Scheduling

  8. Architectural Implications of Agentic AI Workflows

    Aug 5, 2026Jirong Yang, Peizhe Liu, Chaojie Zhang +1LLM Inference SchedulingAgentic Workflows

  9. Learning-Based Collaborative MEC for LLM Inference with Soft-Deadline Awareness via Transformer-Enhanced PPO

    Aug 3, 2026Ngoc Hung Nguyen, Bjorn LandfeldtReinforcement LearningTransformer-Based RL

  10. Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling

    Aug 3, 2026Cunchen Hu, Liangliang Xu, Tian Liu +9Energy-Efficient MLDisaggregated LLM Serving

  11. DeltaServe: Host-Agnostic Co-Serving of Inference and Fine-Tuning for LLMs

    Jul 30, 2026Jiaxuan Chen, Jianshu She, Ye Yuan +5LLM ServingLLM Fine-Tuning

  12. Agentic Coding in the Wild: Characterizing GitHub Copilot Traces at Production Scale

    Jul 30, 2026Banruo Liu, Haoran Qiu, Íñigo Goiri +3LLM InferenceAI Coding Agents

  13. Incast-Free MoE Rate-Based Scheduling

    Jul 28, 2026Evyatar Cohen, Jose Yallouz, Alexander Shpiner +3LLM Inference SchedulingMixture-of-Experts Inference

  14. Conformal Cascade: Distribution-Free Accuracy Guarantees for Multi-Tier LLM Inference

    Jul 27, 2026Yifan Dou, Shikan Lian, Shibo LiLLM Inference EfficiencyLLM Inference Scheduling

  15. Kalypso: Relational LLM Serving

    Jul 26, 2026Hojae Son, Md Ashraful Islam, Huy Gia Cao +2LLM ServingLLM Inference Scheduling

  16. An MLIR-Based Compilation Method for Large Language Models

    Jul 17, 2026Pengchao Hu, Zhibin Xin, Yifan Chen +3LLM ServingLLM Inference Scheduling

  17. Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices

    Jul 11, 2026Yangyijian Liu, Hongyi Ye, Mingyang Li +1On-Device Language Model InferenceLLM Inference Scheduling

  18. SMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling

    Jul 9, 2026Jiahao Wang, Kaizhan Lin, Kaixi Zhang +7LLM Inference SchedulingKV-Cache Management

  19. Sangam: Efficiently Serving Diffusion LLMs with the AR Stack

    Jul 5, 2026Nitin Kedia, Saurabh Agarwal, Myungjin Lee +1LLM Inference SchedulingLLM Inference Serving

  20. Scheduling Thoughts: Learning the Order of Thought in Diffusion Language Models

    Jun 22, 2026Jiawei Xu, Minghui Liu, Aakriti Agrawal +2Masked Diffusion ModelsLanguage Model Decoding

  21. Geometry-Aware Online Scheduling for LLM Serving: From Theoretical Bound to System Practice

    Jun 21, 2026Li Kong, Qi Qi, Yinyu Ye +1LLM ServingLLM Inference Scheduling

  22. Beyond Prediction: Tail-Aware Scheduling for LLM Inference

    Jun 16, 2026Yueying Li, Yuanfan Chen, Jiayang Chen +6LLM ServingInference-Time Optimization