LLM Request Scheduling

LLM: Large Language Model

Momentum

1 paper in the last four weeks, with none the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 16

All topics
CardsList
  1. Denoising Surface: Modeling and Predicting Inference Cost for Diffusion LLM Serving

    Sep 30, 2026Haoyu Zheng, Fangcheng Fu, Binhang Yuan +6Cost-Aware InferenceDiffusion Model Serving

  2. LLM Inference Under Bursty Workload Distribution: Modifying the WAIT Algorithm

    Aug 6, 2026Anjali Gangadhar Katageria, Shobha Rani, Raghu Nandan SenguptaLLM Inference SchedulingOnline Resource Allocation

  3. Sangam: Efficiently Serving Diffusion LLMs with the AR Stack

    Jul 5, 2026Nitin Kedia, Saurabh Agarwal, Myungjin Lee +1LLM Inference SchedulingLLM Inference Serving

  4. Towards Load-Aware Prefill Deflection for Disaggregated LLM Serving

    Jul 2, 2026Shrikara Arun, Anjaly Parayil, Srikant Bharadwaj +2LLM ServingDisaggregated LLM Serving

  5. Geometry-Aware Online Scheduling for LLM Serving: From Theoretical Bound to System Practice

    Jun 21, 2026Li Kong, Qi Qi, Yinyu Ye +1LLM ServingLLM Inference Scheduling

  6. Beyond Prediction: Tail-Aware Scheduling for LLM Inference

    Jun 16, 2026Yueying Li, Yuanfan Chen, Jiayang Chen +6LLM ServingInference-Time Optimization

  7. Harmonizing Real-Time Constraints and Long-Horizon Reasoning: An Asynchronous Agentic Framework for Dynamic Scheduling

    May 28, 2026Shijie Cao, Yuan Yuan, Jing LiuLanguage Model-Based ControlReal-Time Scheduling

  8. PRISM: Fast Online LLM Serving via Scheduling-Memory Co-design

    May 9, 2026Xingyu Qu, Tianhao Lin, Yiqi Li +2LLM ServingKV Caching

  9. Joint Optimization of Trajectory Control, Resource Allocation, and Task Offloading for Multi-UAV-Assisted IoV

    May 6, 2026Maoxin Ji, Qiong Wu, Pingyi Fan +4Trajectory OptimizationIntelligent Transportation Systems

  10. HiveMind: OS-Inspired Scheduling for Concurrent LLM Agent Workloads

    Apr 18, 2026Justice Owusu Agyemang, Jerry John Kponyo, Obed Kwasi Somuah +3LLM Inference EfficiencyAI Coding Agents

  11. Learning to Route and Schedule LLMs from User Retrials via Contextual Queueing Bandits

    Feb 2, 2026Seoungbin Bae, Junyoung Son, Dabeen LeeLLM ServingLLM Routing

  12. Ranking Before Serving: Low-Latency LLM Serving via Pairwise Learning-to-Rank

    Sep 25, 2025Yiheng Tao, Yihe Zhang, Matthew Dearing +4LLM ServingLearning to Rank

  13. LLM Serving Optimization with Variable Prefill and Decode Lengths

    Aug 8, 2025Meixuan Wang, Yinyu Ye, Zijie ZhouLLM ServingLLM Inference Scheduling