Distributed Inference

Latest papers 55

All topics
CardsList
  1. UltraEP: Unleash MoE Training and Inference on Rack-Scale Nodes with Near-Optimal Load Balancing

    Jun 2, 2026Xinming Wei, Chao Jin, Tuo Dai +10Expert ParallelismExpert Load Balancing

  2. NetKV: Network-Aware Decode Instance Selection for Disaggregated LLM Inference

    Jun 2, 2026Mubarak Adetunji OjewaleLLM InferenceKV Caching

  3. Beyond Task-Agnostic: Task-Aware Grouping for Communication-Efficient Multi-Task MoE Inference

    May 31, 2026Zhiyao Xu, Aoxue Liu, Zhanjie Ding +3Mixture-of-Experts InferenceMixture-of-Experts Models

  4. Lodestar: An Online-Learning LLM Inference Router

    May 31, 2026Gangmuk Lim, Wanyu Zhao, Brighten Godfrey +3Inference-Time OptimizationLLM Inference Scheduling

  5. How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving

    May 27, 2026Hanjiang Wu, Abhimanyu Rajeshkumar Bambhaniya, Sarbartha Banerjee +9LLM ServingLLM Inference

  6. Unextractable Protocol Models: Collaborative Training and Inference without Weight Materialization

    May 22, 2026Alexander Long, Chamin Hewa Koneputugodage, Thalaiyasingam Ajanthan +5LLM InferenceDecentralized Learning

  7. Frontier: Towards Comprehensive and Accurate LLM Inference Simulation

    May 20, 2026Yicheng Feng, Xin Tan, Yangtao Deng +3LLM ServingDisaggregated LLM Serving

  8. Federated Martingale Posterior Samping

    May 18, 2026Boning Zhang, Matteo Zecchin, Mingzhao Guo +2Bayesian Neural NetworksPosterior Sampling

  9. Byzantine-Robust Distributed Sparse Learning Revisited

    May 13, 2026Yuxuan Wang, Lixin Zhang, Kangqiang LiRobust RegressionDistributed Optimization

  10. ChunkFlow: Communication-Aware Chunked Prefetching for Layerwise Offloading in Distributed Diffusion Transformer Inference

    May 11, 2026Han Meng, Danny Willow Liu, Dong LiDiffusion TransformerGPU Acceleration

  11. ShardTensor: Domain Parallelism for Scientific Machine Learning

    May 11, 2026Corey Adams, Peter Harrington, Akshay Subramaniam +4Scientific MLDistributed Training

  12. PoHAR: Understanding Hyperlocal Human Activities with Pollution Sensor Networks

    May 10, 2026Prasenjit Karmakar, Karthik Reddy, Sandip ChakrabortyHuman Activity RecognitionInternet of Things

  13. Towards Distributed Inference of LLMs on a P2P Network

    May 7, 2026Shabari S Nair, Krishanu SainiLLM InferenceKV Caching

  14. Enabling Federated Inference via Unsupervised Consensus Embedding

    May 7, 2026Yui Hashimoto, Takayuki Nishio, Yuichi Kitagawa +1Ensemble LearningPrivacy-Preserving ML

  15. MoE-Prefill: Zero Redundancy Overheads in MoE Prefill Serving

    May 3, 2026Zhaoyuan Su, Olatunji Ruwase, Karthik Ganesan +5Expert ParallelismLLM Serving

  16. SpaceMoE: Realizing Distributed Mixture-of-Experts Inference over Space Networks

    May 1, 2026Zhanwei Wang, Huiling Yang, Min Sheng +2Mixture-of-Experts Language ModelsMixture-of-Experts Inference

  17. AI Inference as Relocatable Electricity Demand: A Latency-Constrained Energy-Geography Framework

    Apr 30, 2026Xubin Luo, Cheng YangEnergy-Efficient MLCost-Aware Inference

  18. Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns

    Apr 25, 2026Abhimanyu Bambhaniya, Geonhwa Jeong, Jason Park +6Expert Load BalancingMixture-of-Experts Inference

  19. Time, Causality, and Observability Failures in Distributed AI Inference Systems

    Apr 23, 2026Ankur Sharma, Deep Shah, David Lariviere +1Temporal ConsistencyDistributed Inference

  20. NeuroMesh: A Unified Neural Inference Framework for Decentralized Multi-Robot Collaboration

    Apr 16, 2026Yang Zhou, Yash Shetye, Long Quang +8Efficient Neural Network InferenceMulti-Robot Systems

  21. Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters

    Apr 7, 2025Zonghang Li, Tao Li, Wenjiao Feng +8On-Device Language Model InferenceLLM Inference