cs.AISep 21, 2026

Efficient Iterative Retrieval with Heterogeneous Batching

Authors: Dohyun ParkHubertus FrankeDaniel G. WaddingtonSwaminathan SundararamanYongjoo Park

Abstract

Modern information retrieval increasingly employs both embedding and generative models to handle complex queries. However, current serving systems suffer from low throughput and poor GPU utilization because they execute these models in isolation. Coarse-grained partitioning, such as dedicating GPUs to specific tasks, fails to adapt to dynamic workloads and creates computational "bubbles". To address these, we present Orthrus, a serving system that performs heterogeneous batching within a unified inference loop. The primary challenge lies in unifying embedding and generation workloads with conflicting computational patterns while optimizing batch composition for high performance. Orthrus addresses these challenges through chunked embedding with incremental pooling and by adjusting batch composition in a workload-aware manner. Evaluation on four A100 GPUs shows that, relative to baseline deployments, Orthrus achieves 1.28×\times--4.52×\times higher throughput on controlled workloads and up to 55.8% lower end-to-end p99 latency on an iterative-RAG benchmark. We release our code at https://github.com/illinoisdata/Orthrus .

Explore similar work

Jun 19, 2026cs.DC

BatchGen: An Architecture for Scalable and Efficient Batch Inference

Batch inference has become a central mode of AI computation, yet existing inference engines still rely on execution models designed for interactive serving. When scaled to millions of sequences, batch workloads reveal two fundamental requirements: the ability to handle extreme inter- and intra-sequence load variation that emerges only at runtime, and the ability to sustain high utilization across large fleets of GPUs. Existing systems fail to meet these requirements, losing substantial fractions of achievable throughput. We introduce a new architectural foundation for batch inference: the sequence coroutine compute model, which represents each sequence as a fine-grained, event-driven coroutine. This model exposes expressive primitives that allow the runtime to reorganize work dynamically, enabling larger expert-level batches, mitigating stragglers, reallocating work across devices, and maintaining utilization even on cost-effective or memory-constrained GPUs. Building on this abstraction, we implement BatchGen, a production-ready system that uses the coroutine model at cluster scale. On a 128-GPU cluster, BatchGen reduces batch completion time by up to 2.3×2.3\times, and on memory-constrained accelerators it outperforms the strongest offloading baseline by up to 9.6×9.6\times. We will open-source BatchGen at https://github.com/batchgen-project/batchgen
Tairan Xu, Leyang Xue, Zhan Lu +7
May 30, 2026cs.AI

Threshold-Based Exclusive Batching for LLM Inference

Mixed batching (MB)--interleaving prefill and decode in a single batch--has become the standard scheduling strategy for large language model (LLM) inference due to its efficiency in maximizing compute and memory utilization. However, through controlled experiments, we find that prefill-decode interference inflates MB's per-step marginal cost above that of pure decode. On the high-bandwidth H200 (4.8 TB/s), this occurs only when decode tokens exceed 80% of the batch; however, on the bandwidth-constrained RTX PRO 6000 (1.792 TB/s), this threshold plummets to just 20%. Consequently, the optimal choice between MB and exclusive batching (EB) fundamentally depends on GPU memory bandwidth, model size, and workload composition. We derive a closed-form condition for this EB-MB performance crossover, along with asymptotically optimal phase-switching thresholds and memory-safe batch sizing for EB. Optimized EB achieves up to 41.9% higher throughput on bandwidth-constrained GPUs, while MB retains its advantage on high-bandwidth hardware with larger models. Our hybrid scheduler EB+ applies this condition online to dynamically switch between EB and MB without manual intervention. Under non-stationary traffic with distribution or concurrency shifts, EB+ attains the highest or near-highest throughput in every setting, outperforming MB by up to 36.4%.
Weifang Zhang, Yuzhou Nie, Bowen Pang +2
May 1, 2026cs.DC

SURGE: SuperBatch Unified Resource-efficient GPU Encoding for Heterogeneous Partitioned Data

We present SURGE, a streaming GPU encoding system deployed in production to generate embeddings for over 800 million texts across 40,000 logical partitions. Production embedding pipelines face a tension between logical data partitioning and efficient GPU utilization: processing each partition independently incurs PP inter-process communication (IPC) calls whose overhead limits throughput for compute-light models. Our contributions are analytical: (i) a cost model (Theorem 1) predicting throughput within 2% across three encoders spanning a 15×\times parameter range; (ii) a memory-safety bound (Lemma 3) enabling a streaming two-threshold policy with peak memory O(Bmin+nmax)O(B_{\min} + n_{\max}) rather than O(N)O(N); and (iii) a φφ/CV decision framework characterizing when the pattern applies beyond our workload. The naive fix of batching at fixed size requires O(N)O(N) peak memory (32.7 GB at 10M texts; infeasible beyond ~60M on 192 GB nodes), produces no output until all encoding completes, and offers no fault tolerance. SURGE achieves the same throughput with O(Bmin+nmax)O(B_{\min} + n_{\max}) bounded memory (2.6 GB), 68×\times faster time-to-first-output, and crash recovery at SuperBatch granularity. On 10M texts with 4 NVIDIA L4 GPUs, SURGE delivers 26,413 texts/s -- matching fixed-batch throughput while using 12.6×\times less memory. We validate on bge-base (109M, dd=768, error 1.3%) and across log-normal σσ in {1.0, 1.72, 2.5} (speedup invariant within ±\pm3%), and compare against a partition-batched baseline (PB-PBP-LB), against which SURGE retains a 7% throughput edge and 2.5×\times faster TTFO. Complementary engineering -- zero-copy Arrow serialization (22-25×\times speedup) and async I/O pipelining (up to 93% benefit) -- realizes the design but is not the contribution.
Shashank Kapadia, Deep Narayan Mishra, Sujal Reddy Alugubelli +3