Organizations: Key Lab of High Confidence Software Technologies (Peking University), Beijing, China · State Key Laboratory of Networking and Switching Technology (BUPT), Beijing, China
On-device streaming omni-modal inference safeguards user privacy and eliminates prohibitive per-token API costs, but faces a critical bottleneck: the continuous influx of multimodal data rapidly exhausts constrained memory and compute budgets via monotonic KV cache growth. Existing sparse attention methods fall short, either incurring prohibitive online estimation latency or destroying interleaved cross-modal context, while failing to resolve physical memory fragmentation. We present OmniTide, the first algorithm-system co-design tailored for efficient on-device streaming omni-modal inference. Driven by the observation of modality-aware structural sparsity, OmniTide adopts a unit-based abstraction with two components: (1) At the algorithm level, OmniPick logically retains critical multimodal context based on unit boundaries and modality importance to preserve task accuracy; (2) At the system level, OmniPage physically partitions the cache by retention likelihood and dynamically compacts surviving sparse tokens, minimizing both memory fragmentation and data-movement overhead. Extensive evaluations across three streaming benchmarks and two consumer-device architectures show that OmniTide achieves up to 12.72× kernel speedups and 2.40× lower stream-loop latency. On StreamingBench, it improves accuracy by up to 18.0 percentage points over sliding-window baselines at comparable session cost. OmniPage further reduces the physical KV span by up to 26.7% relative to native logical eviction, unlocking real-time, infinite-context streaming on edge devices.
Figures & tables
Figure 1. Streaming omni-modal inference with modality-specific encoders and an LLM backbone. Explicit unit delimiters are illustrated using MiniCPM-o-4.5 as an example. A streaming omni-modal pipeline showing modality-specific encoders, chunk-sized units, and an omni transformer.
Model
Image 1344×1344
Audio tokens/s
Est. tokens/unit
Est. tokens/min
MiniCPM-o-2.6
640
25
≈700
≈42K
MiniCPM-o-4.5
640
10
≈700
≈42K
Qwen2.5-Omni-7B
2,304
25
≈2,300
≈140K
Table 1. LLM input token counts and estimated audiovisual token growth at one unit per second ( OpenBMB, 2025 ; Cui et al., 2026 ; Xu et al., 2025 ) .
Figure 2. Memory and latency growth with full-history retention on MiniCPM-o-4.5. The model runs on Apple M2 Pro using Q4_K_M LLM weights and FP16 vision/audio projectors, with Metal backend and FlashAttention enabled. Around 24 GB of unified memory is available for inference. Three plots showing inference memory reaching approximately 23 GiB, per-chunk prefill latency rising from 28 to 69 seconds, and attention time increasing from 2 to 39 seconds as the KV cache approaches 80,000 tokens.
Figure 3. History-selection strategies: (a) dense attention, (b) H 2 O heavy-hitter retention, (c) StreamingLLM’s global sink prefix and recent window, and (d) OmniTide’s unit-aware local sinks and recent context. Four schematic attention patterns comparing dense history, H2O heavy hitters, StreamingLLM, and OmniTide with explicit unit boundaries.
Figure 4. OmniTide runtime overview. The frontend assembles temporally ordered multimodal units, OmniPick selects the logically retained history, and OmniPage maintains the physical KV layout used by the transformer loop. A system overview showing the omni frontend creating multimodal units, the transformer and KV cache read-update loop, OmniPick updating logical retention, OmniPage managing physical KV placement, and text or optional audio output.
Figure 5. Selected attention traces from MiniCPM-o-4.5, MiniCPM-o-2.6, and Qwen2.5-Omni (top to bottom). Three rows of attention heatmaps for MiniCPM-o-4.5, MiniCPM-o-2.6, and Qwen2.5-Omni, highlighting local sinks, recent-context concentration, and modality-asymmetric attention.
Figure 6. OmniPick retains a recent window of complete units and selected sink spans from older units. FIFO in this schematic denotes token-level sliding. A schematic of OmniPick retaining recent complete units and selected partial spans, with FIFO shown as a positional comparison.
Figure 7. Logical eviction leaves masked holes between surviving KV entries. Sparse holes keep the exposed physical extent large. Physical KV cells before and after retention, showing masked holes and active or skipped tiles within the exposed physical extent.
Figure 8. Native and eager layouts under OmniPick retention in a 60-chunk MiniCPM-o-4.5 Metal pressure run. Physical extent, occupied cells, and holes for native and eager layouts, alongside their exposed and useful byte footprints.
Figure 9. Class-aware OmniPage placement and selective migration, illustrated with 64-cell pages and 32-cell subblocks. Sparse KV cells organized by allocation class, with retained cells selectively relocated to reclaim pages within the existing tensor.
Figure 10. CUDA latency at 4K–32K no-slide checkpoints. Panels show three models with four checkpoints each. CUDA streaming-update and attention-kernel latency for MiniCPM-o-4.5 and Qwen2.5-Omni-3B/7B at four history lengths.
Figure 11. Metal latency at matched no-slide checkpoints, with normal update timing above separate attention profiling. Metal streaming-update and attention-kernel latency for three models at 4K, 8K, 16K, and 32K no-slide histories.
Qwen2.5-Omni
MCPM-o-4.5
Method
3B
7B
9B
No-slide
62.3 / 43.4 / 30.2
52.0 / 55.1 / 31.2
74.1 / 43.4 / 17.7
FIFO / basic
58.1 / 42.9 / 29.2
49.8 / 53.9 / 30.2
69.0 / 40.6 / 17.3
Token-basic
56.2 / 42.2 / 28.9
47.8 / 52.6 / 30.3
57.0 / 40.1 / 16.8
StreamingLLM
59.6 / 42.6 / 28.2
50.5 / 47.8 / 31.4
70.1 / 41.3 / 17.5
H2O
55.8 / 40.1 / 28.2
47.0 / 49.2 / 29.2
52.7 / 35.1 / 17.0
Table 2. Task-quality summary. Each cell reports StreamingBench / SVBench / LiveSports-3K-cc . StreamingBench reports percentages; SVBench reports 0–100 Terra dialogue scores; LiveSports reports 0–100 Terra semantic-alignment scores.
Figure 12. CUDA KV layout at the 32K checkpoint. Stacked bars show useful cells and physical-span waste (left axis); the line shows cumulative peak physical span (right axis). Labels above bars give span/useful-cell ratios. Three model panels comparing useful KV, physical span, and peak span for StreamingLLM, OmniPick only, H2O, and OmniTide.
Figure 13. StreamingBench quality–efficiency tradeoff. Lower left is better: session time decreases to the left and accuracy increases downward. Dashed horizontal lines mark no-slide accuracy; N/A bands denote unreported metrics. Three panels for Qwen2.5-Omni-3B, Qwen2.5-Omni-7B, and MiniCPM-o-4.5 comparing mean stream-loop time and accuracy. H2O appears in the latency N/A band; PagedAttention and vAttention appear in the accuracy N/A band. OmniTide combines low session time with high accuracy among the plotted configurations.
Model
No-slide (s)
OmniTide (s)
Speedup
Qwen2.5-Omni-3B
40.33
17.61
2.290 ×
Qwen2.5-Omni-7B
51.71
21.52
2.403 ×
Table 3. End-to-end session latency on CUDA, excluding model initialization and system-prompt prefill.
Figure 14. Watermark sensitivity on StreamingBench. The vertical axis is accuracy (%); horizontal labels give high/low watermarks, with no-slide as a reference. Three accuracy curves across high/low watermark settings, with no-slide references.
Figure 15. Cumulative attention GPU time under identical retention. OmniPage includes migration and synchronization costs; Eager excludes them. Prefill and decode cumulative FlashAttention GPU times, with OmniPage reductions of 26.1 and 22.5 percent relative to native layout.
On-device LLM inference is attractive for privacy and responsiveness, but remains challenging on mobile and embedded devices because model weights far exceed available DRAM. Prior systems exploit activation sparsity and offload weights to SSD or flash storage, but face a fundamental systems trade-off: accurate sparse execution decisions require the latest context, whereas efficient computation-I/O overlap requires early prediction. As a result, existing designs either serialize execution or incur redundant weight fetches, extra computation, and large cache overheads. We present LeanStream, a streaming speculate-and-refine framework for efficient on-device LLM inference. LeanStream progressively refines computation, loading, and cache-retention priorities using partial GPU results, enabling fine-grained overlap between GPU execution and storage I/O. We implement LeanStream on both mobile and embedded platforms. Compared with prior on-device LLM inference systems, LeanStream reduces memory usage by 4.8× to 7.5× at the best throughput achieved by prior work, while further improving token generation throughput by 1.6× to 2.1×.
Renyuan Liu, Yuyang Leng, Kaiyan Liu +7
Richard · George Mason University · Global Technology Applied Research, JPMorganChase +1
Efficient inference for on-device Large Language Models (LLMs) remains challenging due to limited hardware resources and the high cost of the prefill stage, which processes the full input context to construct Key-Value (KV) caches. We present SparKV, an adaptive KV loading framework that combines cloud-based KV streaming with on-device computation. SparKV models the cost of individual KV chunks and decides whether each chunk should be streamed or computed locally, while overlapping the two execution paths to reduce latency. To handle fluctuations in wireless connectivity and edge resource availability, SparKV further refines offline-generated schedules at runtime to rebalance communication and computation costs. Experiments across diverse datasets, LLMs, and edge devices show that SparKV reduces Time-to-First-Token by 1.3$x-5.1x with negligible impact on response quality, while lowering per-request energy consumption by 1.5x to 3.3x, demonstrating its robustness and practicality for real-world on-device deployment.
Hongyao Liu, Liuqun Zhai, Junyi Wang +1
Department of Computer Science, City University of Hong Kong, Hong Kong SAR.
On-device mobile Large Language Model (LLM) inference is gaining significant attention. However, mobile devices operate in highly dynamic multitasking environments where users frequently switch between applications. This creates memory pressure, forcing LLM memory (model weights and KV cache) to be evicted by the operating system. When a new inference request arrives, the inference system must restore the evicted memory through slow storage reads or recompute the entire KV cache, severely degrading responsiveness. To address this, we present mzCache, an on-device LLM inference system with specialized memory management for multitasking environments. Under unpredictable memory pressure, mzCache elastically evicts LLM memory and leverages the unified memory of mobile SoCs to enable zero-wait inference on the GPU with concurrent CPU-side restoration. mzCache realizes this through restoration-oriented memory management: LLM memory is partitioned into fine-grained shared buffers to enable partial eviction and restoration with concurrent cross-processor access, while hybrid swap and backward-out eviction policies ensure low-latency restoration from any eviction state. Implemented on llama.cpp and deployed as an Android application, mzCache achieves 2.1-5.5× reduction in Time-to-First-Token compared to storage-backed partial offload and demonstrates its effectiveness in real multitasking scenarios.