LLM Inference Acceleration

LLM: Large Language Model

Latest papers 677

All topics
CardsList
  1. PReCache: Efficient KV Cache Sharing for Multi-LoRA Agents via Low-Rank Precomputation and Neutral Reconstruction

    Sep 28, 2026Hyesung Jeon, Hyeongju Ha, Jae-Joon KimMulti-Agent LLM SystemsMemory-Efficient Inference

  2. Where Activation Sparsity and KV-Cache Sparsity Cross in LLM Decoding

    Sep 27, 2026Jungseob Lee, Seungyoon Lee, Seongtae Hong +2Language Model DecodingActivation Sparsity

  3. Faster Block-Diffusion Serving with Distribution-Free Risk Guarantees

    Sep 27, 2026Jungseob Lee, Dongyub Jude Lee, Chanjun Park +2Diffusion Language Model InferenceLLM Inference Acceleration

  4. PQ-HSA: Reusing Product-Quantized Scores for Hybrid Sparse-Approximate Attention

    Sep 27, 2026Kunming Shao, Jierun Chen, Yanli Wang +4Self-AttentionLLM Inference Acceleration

  5. Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration

    Sep 27, 2026Shangzhen Zhu, Muyan Hu, Tomasz KozlowskiEfficient Transformer InferenceSoftmax Attention

  6. Does Execution Require Target KV Fidelity? A Mixed-Fidelity KV Runtime for LLM Serving

    Sep 27, 2026Jiantong Jiang, Yue Yang, Peiyu Yang +1LLM ServingKV-Cache Management

  7. DISCO: Distributed Long Context Scaling with Grounding-Reasoning Disaggregation

    Sep 27, 2026Guanzheng Chen, Viet Dac Lai, Subhojyoti Mukherjee +5LLM GroundingLong-Context Language Model Inference

  8. Just Let Linear States Forget the Distant Past: Prefix Caching via Suffix Replay for Hybrid LLMs

    Sep 27, 2026Yirui Liu, Ruoling Qi, Xuaner Wu +6Long-Context Language Model InferenceLLM Inference Acceleration

  9. OLED-MoE: Accelerating MoE-Based dLLM Inference via Inter-Iteration Locality-Aware Expert Offloading

    Sep 27, 2026Jingyuan Xiao, Jiayue Wang, Yitao Hu +7Expert OffloadingMixture of Experts

  10. Multi-Dimensional Comparative Scale Construction for Efficient Personalized Subjective Judgment in High-Traffic Applications

    Sep 27, 2026Xianglong Shi, Shifeng Liu, Sirui Zhao +2Human Preference EvaluationLearning to Rank

  11. SketchSSM: Write to the Full State, Read from a Compact Sketch

    Sep 27, 2026Omin Kwon, JoongWon Shin, Minseo Kim +3LLM Inference AccelerationLinear Attention

  12. Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs

    Sep 26, 2026Vincent-Daniel Yun, Woosang Lim, Haneul Yoo +3Multi-Agent LLM SystemsLLM Inference Acceleration

  13. MILO: Efficient Many-shot In-Context Learning with Block-wise Low-rank Compression

    Sep 24, 2026Youpeng Zhao, Tian Tan, Liqian Peng +2Memory-Efficient InferenceLLM Inference Acceleration

  14. FLEET: From Logits Entropy to Enhanced Trajectories in Text Generation

    Sep 23, 2026Oleksii Streltsov, Oleksandra VitkoLanguage Model DecodingMemory-Augmented Language Models

  15. When Parallel Drafter Meets Parallel Speculative Decoding

    Sep 23, 2026Fuliang Liu, Xue Li, Kun Qian +4Speculative DecodingLLM Inference Acceleration

  16. Distilling Sequential Computation in Transformer Language Models

    Sep 23, 2026Zixuan Lan, Jessica Yang, Yanhong Li +2LLM CompressionLong-Context Language Model Inference

  17. HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing

    Sep 22, 2026Jianyu Wei, Yizhao Gao, Qihao Zhang +12Long-Context RetrievalKV-Cache Management

  18. Disaggregated Quantization: Specializing LLM Prefill and Decode

    Sep 22, 2026Andrei Panferov, Maximilian Kleinegger, Sweta Priyadarshi +2LLM QuantizationLLM Inference

  19. CompKV: Compensation-Aware KV Selection for Long-Context LLM Inference

    Sep 22, 2026Zhen Huang, Ruizhe Yao, Danyi Liu +8Long-Context Language Model InferenceKV-Cache Management

  20. Latest Exact Match Attention

    Sep 22, 2026Moritz BrösamleTransformer ExpressivityMemory-Augmented Language Models

  21. Accelerating the Mitigation of LLM Inference Nondeterminism Across GPU Architectures

    Sep 22, 2026Liam Cooper, Shinnung Jeong, Hyeran Jeon +2GPU Kernel OptimizationLLM Inference Acceleration

  22. ARM: Attention with Routed-Memory for Learnable Sparse Control

    Sep 21, 2026Qiuhao Zeng, Jerry Huang, Peng Lu +7Memory-Augmented Language ModelsLong-Context Language Model Inference

  23. H-Spec: Parallel Speculative Decoding Without a Drafter-Side KV Cache

    Sep 21, 2026Weifan Jiang, Krishna Teja Chitty-Venkata, Megan Flynn +6KV-Cache ManagementSpeculative Decoding