Offloading

Momentum

8 papers in the last four weeks, against 2 the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 39

All topics
CardsList
  1. Task-Oriented Communications for Edge-Assisted Multi-View Localization

    Sep 28, 2026Zhengru Fang, Huanhuan Lou, Senkang Hu +5Drone-View Geo-LocalizationMulti-Uav

  2. MaskCoFT: Masked Co-Adaptive Fine-Tuning for Memory-Efficient MoE Inference

    Sep 28, 2026Junfeng Wu, Zehao Fan, Hadjer Benmeziane +3Mixture-Of-ExpertsParameter-Efficient Fine-Tuning Methods

  3. EfficientAgent: What Makes KV Cache Offloading Work for Concurrent Agents?

    Sep 27, 2026Kunming Shao, Jierun Chen, Jiangnan Yu +7Kv-Cache ManagementOffloading

  4. OLED-MoE: Accelerating MoE-Based dLLM Inference via Inter-Iteration Locality-Aware Expert Offloading

    Sep 27, 2026Jingyuan Xiao, Jiayue Wang, Yitao Hu +7Mixture-Of-Expert InferenceOffloading

  5. TierKV: Long-Context On-Device LLMs via Predictive Multi-Tier KV Caching

    Sep 18, 2026Zhihao Shu, Md Musfiqur Rahman Sanim, Jie Hu +4Kv-Cache ManagementDepthweave-Kv

  6. Designing Against Deskilling: Metacognitive Feedback Reduces Cognitive Offloading to LLM Assistants

    Sep 17, 2026Sebastian Maier, Kai Schwabe, Manuel Schneider +1MetacognitionOffloading

  7. The Operable Pareto Front: Distilling Offline Search into Run-Time Control for Multi-Objective UAV Edge-Computing Scheduling

    Sep 16, 2026Qiao Liao, Zhiyong Feng, Bin Wu +1SchedulersEdge Devices

  8. SeqMoE: Toward Full-Load Performance via Predictive and Graph-Compatible MoE Offloading

    Sep 14, 2026Zihan Wang, Yuqi Wang, Lei Gong +5Mixture-Of-Expert InferenceOffloading

  9. SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations Center

    Sep 3, 2026Uday Vallabhaneni, Cassie L. Cagwin, David J. WildLarge Language Model AgentsCybersecurity

  10. Update for Decisions, Not Freshness: Goal-Oriented Status Updating and Selective Offloading at the Network Edge

    Sep 1, 2026Jianpeng Qi, Qiyang Zhang, Chao Liu +5Edge PlatformsOffloading

  11. Active Beyond-Diagonal RIS Empowered Heterogeneous Edge Computing: A Distributional Reinforcement Learning Approach

    Jul 14, 2026Tianyu Pang, Hongyu LiPower AllocationReconfigurable Intelligent Surfaces

  12. Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices

    Jul 11, 2026Yangyijian Liu, Hongyi Ye, Mingyang Li +1Cpu-Gpu Hybrid DesignsLLM Inference Optimization

  13. Is Your NPU Ready for LLMs? Dissecting the Hidden Efficiency Bottlenecks in Mobile LLM Inference

    Jul 6, 2026Guanyu Cai, Ruiming Tian, Lang Yang +4Neural Processing UnitsLLM Inference Optimization

  14. Extending Responsibility-Sensitive Safety for the Assessment of Offloaded Autonomous Driving Services

    Jun 5, 2026Robin Dehler, Aryan Thakur, Michael BuchholzAutonomous DrivingOffloading

  15. Offloading Score: Measuring AI Reliance Through Counterfactual Workflows

    May 28, 2026Vishakh Padmakumar, Lujain Ibrahim, Zora Zhiruo Wang +3Cognitive LoadOffloading

  16. Knowledge Offloading: Decomposing LLMs into Sparse Backbones and Memory Modules

    May 27, 2026Karim Galliamov, Rochelle Choenni, Ivan TitovOffloadingLow-Rank Adaptation Adapters

  17. TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload

    May 19, 2026Zhiben Chen, Youpeng Zhao, Yang Sui +2Efficient InferenceDiffusion Language Models

  18. DAG-Based QoS-Aware Dynamic Task Placement for Networked Multi-Stage Control Pipelines

    May 19, 2026Thien Tran, Jonathan Kua, Thuong Hoang +3Perception PipelineOffloading

  19. Towards Multi-Model LLM Schedulers: Empirical Insights into Offloading and Preemption

    May 19, 2026Mert Yildiz, Pietro Spadaccino, Alexey Rolich +2PreemptionOffloading

  20. Heterogeneous Tasks Offloading in Vehicular Edge Computing: A Federated Meta Deep Reinforcement Learning Approach

    May 18, 2026Yaorong Huang, Jingtao Luo, Xuechao WangEdge DevicesOffloading

  21. ChunkFlow: Communication-Aware Chunked Prefetching for Layerwise Offloading in Distributed Diffusion Transformer Inference

    May 11, 2026Han Meng, Danny Willow Liu, Dong LiOffloadingPreemption

  22. GELATO: Generative Entropy- and Lyapunov-based Adaptive Token Offloading for Device-Edge Speculative LLM Inference

    May 11, 2026Zengzipeng Tang, Yuxuan Sun, Wei Chen +2LLM Inference OptimizationEdge Devices

  23. ORICF -- Open Robotics Inference and Control Framework

    May 10, 2026Andrés Meseguer Valenzuela, Luís Miguel Bartolín ArnauFast InferenceOffloading

  24. Adaptive DNN Partitioning and Offloading in Heterogeneous Edge-Cloud Continuum

    May 10, 2026Akuen Akoi Deng, Eimantas Butkus, Alfreds Lapkovskis +1Edge DevicesEdge Platforms

  25. Hard to Read, Easy to Jailbreak: How Visual Degradation Bypasses MLLM Safety Alignment

    May 8, 2026Zhixue Song, Boyan Han, Yiwei Wang +1Recent Vision-Language ModelsMultimodal Large Language Models

  26. Joint Optimization of Trajectory Control, Resource Allocation, and Task Offloading for Multi-UAV-Assisted IoV

    May 6, 2026Maoxin Ji, Qiong Wu, Pingyi Fan +4Multi-UavPower Allocation

  27. DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference

    Apr 29, 2026Bodon Jeong, Hongsu Byun, Youngjae Kim +4LLM Inference OptimizationOffloading

  28. Privatar: Scalable Privacy-preserving Multi-user VR via Secure Offloading

    Apr 19, 2026Jianming Tong, Hanshen Xiao, Krishna Kumar Nair +5Virtual RealityPrivacy-Preserving Framework

  29. TrigReason: Trigger-Based Collaboration between Small and Large Reasoning Models

    Apr 16, 2026Yi Zhao, Yajuan Peng, Cam-Tu Nguyen +4Large Reasoning ModelsInference Latency

  30. Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters

    Apr 7, 2025Zonghang Li, Tao Li, Wenjiao Feng +8LLM Inference OptimizationOffloading

  31. FluxMoE: Decoupling Expert Residency for High-Performance MoE Serving

    Date pendingQingxiu Liu, Yongchao He, Runhan Jiang +4Mixture-Of-Expert InferenceMixture-Of-Experts