As large language models (LLMs) keep growing in size and complexity, their training frameworks evolve at a rapid pace as well. Therefore, continuous integration (CI) is critical for maintaining the quality and stability of these frameworks. However, unlike traditional software, CI for LLM training frameworks relies on GPU-intensive tests, which usually involve complete model training or evaluation. This leads CI itself to become a new bottleneck for fast-paced development. In this paper, we introduce FastCI, a framework that improves the efficiency of CI for LLM training frameworks. FastCI leverages runtime evidence to select affected tests and prune tests that execute changed code in equivalent contexts. Then FastCI prioritizes high-risk tests to expose potential failures earlier, and optimizes test workloads along dimensions outside the intended validation scope of each test. Evaluated on the CI workload of our LLM training framework, FastCI reduces the CI latency by 77.5% and the GPU resource usage by 63.9%, while improving the modified code coverage retention by 3.2%, compared with the currently deployed CI pipelines. FastCI has now been integrated into the CI pipelines of our LLM training framework at ByteDance.
Figures & tables
Figure 1: Test selection using path mappings and runtime evidence.
Figure 2: Overview of FastCI.
Figure 3: Affected-test selection for two MRs, either directly from runtime evidence or through within-class expansion when evidence is missing.
Figure 4: Context-aware test pruning when materialize_step_inputs is modified.
Figure 5: End-to-end CI performance. The statistics are relative to full CI.
Figure 6: CDF of CI latency and queueing time under different approaches. The x-axis is shown up to 360 minutes and annotations report the fraction of full CI runs beyond this range.
Method
Number of selected tests
GPU resource cost
Recall of affected tests
Modified code coverage retention
Path-based
62.6%
67.6%
76.1%
96.0%
NameRTS
77.2%
80.8%
94.0%
95.8%
LLM reasoning
39.4%
42.1%
54.6%
83.1%
FastCI-select w/o pruning
35.9%
39.4%
84.5%
99.8%
FastCI-select
25.7%
28.6%
66.2%
99.2%
Table 1: Comparison between different test selection methods.
Figure 7: Effectiveness of risk-aware test scheduling compared with production order, single-indicator variants, and COLEMAN.
Sampling rate
Tracing overhead
Recall of affected tests
Modified code coverage retention
10 Hz
0.3%
57.7%
97.5%
50 Hz
4.1%
66.2%
99.2%
100 Hz
17.0%
68.4%
99.6%
Table 2: Tracing overhead and final test-selection effectiveness at different sampling rates.