Deep learning training commonly runs on shared multi-tenant GPU servers, where exclusive allocation provides isolation but can leave resources underutilized and increase queueing time. Collocation can improve efficiency, but interference-agnostic placement may cause severe slowdowns, while inaccurate memory information can lead to out-of-memory (OOM) failures. We present AEGIS, a server-scale runtime scheduling system for controlled collocation of deep learning training workloads on shared multi-GPU servers. AEGIS integrates memory feasibility, post-placement observation, runtime-pressure filtering, placement, and OOM-aware recovery in a single scheduling loop. After placement, AEGIS observes workload activity before permitting further collocation, then uses low-overhead telemetry to determine whether a GPU can safely accept additional work. OOM failures trigger retries under progressively safer memory conditions, eventually falling back to exclusive execution. This online approach avoids costly offline pairwise compatibility profiling. We evaluate AEGIS using vision, Transformer, recommendation, and LLM-style workloads across three production-derived traces. AEGIS reduces geometric-mean makespan by 16% relative to Lucid, 21% relative to Horus, and 27% relative to exclusive allocation. Sensitivity studies show that activity-anchored observation and runtime-pressure filtering balance conservative isolation against interference-agnostic collocation, improving makespan while limiting sharing-induced per-task slowdown.
Figures & tables
Figure 1 . Interference-agnostic MPS-enabled collocation on an A100 40 GB GPU. We start with ResNet50 (BS128) and VGG16 (BS64), then add Xception (BS64), followed by ResNet50 (BS32), all on ImageNet ( Deng et al., 2009 ) , where BS denotes batch size. All jobs complete in every configuration. Throughput gain is the sum of the jobs’ solo runtimes divided by the collocated makespan.
Figure 2 . Overview of AEGIS . Numbered markers show the control flow: submit a job ( 1 ), select the next job ( 2 ), read its memory input ( 3 ), read the current per-GPU state ( 4 ), place the job ( 5 ), start a mandatory post-placement observation hold ( 6 ), detect activity and update the activity-anchored monitoring window ( 7 ), release the hold and refresh runtime-pressure eligibility ( 8 ), and recover from an OOM ( 9 ).
Group
Model
Dataset
BS
GPUs
Mem. (GB)
∗
ResNet 18/34/50
ImageNet/CIFAR-100
32–128
1
0.6–21.2
∗
MobileNet
ImageNet/CIFAR-100
32–128
1
2.4–3.3
∗
EfficientNet
ImageNet/CIFAR-100
32–128
1
2.6–4.5
∗
VGG16
ImageNet
32–128
1
6.8–21.2
∗
Xception
ImageNet
32–128
1
6.7–20.7
∗
Inception
ImageNet
32–128
1
5.5–17.0
Table 1 . Summary of workloads used in the evaluation. Mem. reports the observed/profiled peak GPU memory range in GB across evaluated specifications (BS: batch size).
Metric (h)
Trace
Exclusive
Horus
Lucid
AEGIS Est.-Free
Makespan
Philly
11.84
10.84
9.94
8.39
Saturn
13.17
12.01
11.22
9.24
Venus
12.35
11.67
11.17
9.66
Mean JCT
Philly
5.23
4.52
3.93
3.31
Saturn
4.78
4.00
3.56
2.79
Venus
4.55
4.16
3.90
2.83
Table 2 . Absolute end-to-end performance across traces. Values are hours; p95 values are within-trace job percentiles. Lower is better for all metrics; bold marks the best value in each row.
Figure 3 . End-to-end trace performance. Values are normalized to Exclusive on the corresponding trace; lower is better. Bars show per-trace values and GeoMean. In the JCT panel, diamonds show p95 JCT normalized to Exclusive p95 JCT on the same trace. Results include failed-attempt runtime and recovery wait.
Figure 4 . Per-job normalized JCT distributions across traces. Each job’s JCT is divided by its isolated solo runtime; curves farther left are better. The x-axes are logarithmic.
Figure 5 . Queueing and execution-time trade-off. Values are normalized to Exclusive on the corresponding trace; lower is better. Bars show means. Diamonds show p95 values normalized to the corresponding Exclusive p95 for the same metric and trace.
JCT
Queue
Exec.
Mem. input
Mkspn
Mean
P95
Mean
P95
Mean
P95
F/R
PeakMem
0.712
0.594
0.607
0.442
0.520
1.785
2.110
0/0
Est-free
0.730
0.613
0.598
0.494
0.546
1.550
1.971
12/12
AnalyticalMem
0.733
0.624
0.638
0.520
0.584
1.453
1.964
3/3
GPUMemNet
0.737
0.629
0.636
0.510
0.573
1.575
2.057
1/1
FakeTensor
0.794
0.648
0.679
0.543
0.641
1.483
1.833
0/0
Table 3 . Effect of memory input. Values are GeoMeans across traces, normalized to Exclusive; lower is better for the time metrics; F/R reports failed/recovered attempts. Bold marks the best value in each metric column; the italic row label marks the default AEGIS configuration. Queue includes recovery wait; Exec. includes failed-attempt runtime.
JCT
Queue
Exec.
Policy
Mkspn
Mean
P95
Mean
P95
Mean
P95
F/R
RR
0.776
0.635
0.636
0.501
0.573
1.686
1.993
28/28
MU
0.755
0.587
0.619
0.449
0.556
1.683
2.160
13/13
MFM
0.730
0.613
0.598
0.494
0.546
1.550
1.971
12/12
Table 4 . Effect of placement policies. Values are GeoMeans across traces, normalized to Exclusive; lower is better for the time metrics; F/R reports failed/recovered attempts. Bold marks the best value in each metric column; the italic row label marks the default AEGIS policy. Queue includes recovery wait; Exec. includes failed-attempt runtime.
JCT
Queue
Exec.
Config.
Mkspn
Mean
P95
Mean
P95
Mean
P95
F/R
MFM
0.730
0.613
0.598
0.494
0.546
1.550
1.971
12/12
MFM-NF
0.732
0.622
0.646
0.470
0.570
1.821
2.265
18/18
MU
0.755
0.587
0.619
0.449
0.556
1.683
2.160
13/13
MU-NF
0.770
0.616
0.626
0.480
0.580
1.740
2.250
19/19
Table 5 . Runtime-pressure eligibility ablation. Values are GeoMeans across traces, normalized to Exclusive; lower is better for the time metrics; F/R reports failed/recovered attempts. Bold marks the best value in each metric column; the italic row label marks the default. Queue includes recovery wait; Exec. includes failed-attempt runtime. NF: saturation-based pressure filtering disabled; the post-placement observation hold remains enabled.
JCT
Queue
Exec.
Config.
Mkspn
Mean
P95
Mean
P95
Mean
P95
F/R
AEGIS (Mem-only)
1.007
0.998
1.076
0.116
0.155
1.731
1.955
0/0
AEGIS (U70%)
0.884
0.892
0.858
0.585
0.754
1.172
1.115
0/0
Horus
0.897
0.890
0.844
0.647
0.760
1.112
1.044
0/0
Lucid
0.950
0.876
0.832
0.510
0.631
1.195
0.963
0/0
Table 6 . Limited-telemetry evaluation on the heterogeneous RTX/GTX server. Values are GeoMeans across traces, normalized to Exclusive; lower is better for the time metrics; F/R reports failed/recovered attempts. Bold marks the best value in each metric column; the italic row label marks the selected AEGIS fallback. Queue includes recovery wait; Exec. denotes execution time. U70%: 70% utilization threshold.
Figure 6 . Stability of activity-anchored monitoring windows. Points show weighted mean absolute error of the AEGIS pressure indicators relative to a 200 s activity-anchored reference.
Figure 7 . Pressure-threshold trade-off on calibrated one-GPU workload sequences. The top panel shows throughput gain and slowdown; the bottom panel shows collocation deferrals. The selected operating point lies near the knee of the trade-off.
LLM-agent services repeatedly execute small deterministic transitions between model and tool calls: route an outcome, update state, and emit the next effect. We ask when this control path exposes enough concurrent work for GPU execution, and what changes when a GPU-computed route decision remains on device. We formalize the ready-cohort boundary using fixed-partition share F, exact offline share P*, local upper bound U, and online achieved share A. Under zero service time, unlimited capacity, and equal relative launch deadlines, a specialized dynamic program computes P* exactly. In a stationary Poisson replay of one pinned 851-session public trace panel, the primary condition at 100,000 target active sessions, K=256, and a 50 ms launch deadline gives F=30.19%, P*=43.00%, and U=45.85%. Exact packing recovers 81.83% of the opportunity lost at fixed window boundaries. The outcome-derived route key is a conditioning proxy, not proof of executable identity. A separate mechanism study keeps a GPU-computed binary decision on device instead of returning four bytes to the host and redispatching. Across four named GPU placements, the device-resident path is faster in all 36 configurations; within-placement row-median ratios range from 1.19x to 2.39x. Across both admissible mechanisms, all 14,557,440 tested batched invocations match a separately implemented host oracle. A fixed nested device graph that removes no host decision is slower in all 60 configurations across five placements. Together, the studies establish two measurable gates for GPU agent control: deadline-feasible cohort supply and observation placement. A joined finite online runtime is required to measure A, CPU displacement, and service-level benefit.
With the rapid advancement of deep learning technology, shared GPU clusters receive an increasing number of deep learning training (DLT) jobs. Yet resource fragmentation make such clusters underutilized and forces the DLT jobs running on them to endure long turnaround times. Extensive research has been devoted to quantifying fragmentation and developing scheduling algorithms that alleviate its impact. However, existing fragmentation measures break down in the absence of workload distribution information, while current schedulers cannot continuously maintain resource fragmentation at a low level. To tackle these problems, we first introduce Scheduler-Induced Fragmentation (SIF), a metric built on the notion of partial-nodes that is independent of historical workload knowledge. We then propose COMPASS-ABS, which employs the COMPact-ASSured (COMPASS) algorithm to confine the cluster state within a tight Anchor-Based Space (ABS), whose construction fully leverages the topological alignment between dominant workload size and node capacity. Moreover. We also prove that it ensures SIF is bounded by N2 under a workload composition condition that matches both theory and production. Evaluations implemented on a physical cluster and a simulated cluster demonstrate COMPASS-ABS effectiveness at improving resource utilization, reducing DLT job completion time by reducing fragmentation.
AI agents execute tens to hundreds of chained LLM calls per task, yet GPU schedulers treat each call as independent, discarding gigabytes of intermediate state between steps and inflating end-to-end latency by 3-8x. We argue that this request-level abstraction is fundamentally mismatched to compound AI workloads, and propose a shift to program-level scheduling: treating the entire agent workflow (not individual inference calls) as the first-class schedulable unit. We present SAGA, a distributed scheduler that implements this abstraction through three mechanisms: (1) Agent Execution Graphs that capture workflow structure to predict KV cache reuse across tool-call boundaries, achieving within 1.31x of Bélády's optimal offline policy; (2) session-affinity batching with work stealing that co-locates correlated requests while maintaining global load balance; and (3) Agent Fair Share, a task-completion-time fairness metric with provable bounded-deviation guarantees. On a 64-GPU cluster serving SWE-bench coding agents and WebArena browser tasks, SAGA reduces task completion time by 1.64x (geometric mean, p < 0.001) over vLLM v0.15.1 with prefix caching and affinity routing, while improving GPU memory utilization by 1.22x and achieving 99.2% SLO attainment under multi-tenant interference. These latency gains come at a quantified cost: approximately 30% lower peak throughput than throughput-optimal batch scheduling, a tradeoff appropriate for the latency-sensitive interactive deployments that dominate compound AI usage. Our results demonstrate that workflow-aware scheduling is essential for efficient compound AI serving.