cs.LGOct 1, 2026

FastCI: Efficient GPU-Intensive CI for LLM Training Frameworks

Authors: Tianshuo Qiao, Naiqian Zheng, Xiaopeng Liu, Shuguang Wang, Diandian Gu, Xuanzhe Liu, Xin Jin

Organizations: Peking University · ByteDance Seed

Abstract

As large language models (LLMs) keep growing in size and complexity, their training frameworks evolve at a rapid pace as well. Therefore, continuous integration (CI) is critical for maintaining the quality and stability of these frameworks. However, unlike traditional software, CI for LLM training frameworks relies on GPU-intensive tests, which usually involve complete model training or evaluation. This leads CI itself to become a new bottleneck for fast-paced development. In this paper, we introduce FastCI, a framework that improves the efficiency of CI for LLM training frameworks. FastCI leverages runtime evidence to select affected tests and prune tests that execute changed code in equivalent contexts. Then FastCI prioritizes high-risk tests to expose potential failures earlier, and optimizes test workloads along dimensions outside the intended validation scope of each test. Evaluated on the CI workload of our LLM training framework, FastCI reduces the CI latency by 77.5% and the GPU resource usage by 63.9%, while improving the modified code coverage retention by 3.2%, compared with the currently deployed CI pipelines. FastCI has now been integrated into the CI pipelines of our LLM training framework at ByteDance.

Figures & tables

Explore similar work

CardsList
  1. A Few GPUs, A Whole Lotta Scale: Faithful LLM Training Emulation with PrismLLM

    May 15, 2026Shaoke Xi, ChonLam Lao, Boyi Jia +11Large Language Model TrainingLarge Language Models(Llms

  2. FastKernels: Benchmarking GPU Kernel Generation in Production

    May 22, 2026Gabriele Oliaro, Yichao Fu, May Jiang +5Graphics Processing Unit Kernels

  3. Hybrid JIT-CUDA Graph Optimization for Low-Latency Large Language Model Inference

    Apr 25, 2026Divakar Kumar Yadav, Tian ZhaoLLM Inference OptimizationInference Latency