Diffusion-2BC: Hybrid Diffusion and Regression Training for Offline Behavior Cloning in Autonomous Driving
Authors: Bruno Maciel Machado, Eric Aislan Antonelo
Organizations: Department of Automation and Systems Engineering, Graduate Program in Automation and Systems Engineering, Federal University of Santa Catarina, Florianópolis, Santa Catarina, Brazil
Behavior cloning provides an offline route to autonomous-driving policy learning, but mean-squared-error regression is poorly matched to demonstrations in which one observation admits several valid actions. Diffusion policies can represent conditional multimodal action distributions, yet their closed-loop performance may be unstable when visual features and control are learned from limited data. This paper presents Diffusion-2BC, which combines a diffusion denoising objective with an auxiliary deterministic behavior-cloning loss over a shared visual encoder. The auxiliary branch is used only during training; inference remains diffusion-based. The proposed method is evaluated in the controlled Claw environment and in bird's-eye-view CARLA navigation, including route-conditioned driving, route-free navigation through multiple intersections, and cross-map evaluation from Town01 to Town02. In the Claw task, Diffusion-2BC reduced the mean mask-distance error by approximately 10% relative to a diffusion-based behavior-cloning baseline and by 85% relative to standard deterministic behavior cloning. In route-free CARLA, Diffusion-2BC traveled substantially farther before termination under the evaluation protocol than both baselines in Town01 and Town02. Additional qualitative rollouts revealed distinct route choices, showing the multimodal behavior of the proposed diffusion-based agent. The results indicate that an auxiliary regression signal can improve the closed-loop reliability of diffusion behavior cloning while preserving multimodal prediction in the controlled benchmark.
Figures & tables
Method family
Action representation
Handles unlabeled multimodality
Main strength
Limitation relative to this study
MSE behavior cloning
Deterministic point estimate
No
Simple, fast offline training and inference
Conditional averaging or single-mode behavior
Command-conditioned BC
Deterministic action conditioned on a route command
Only after labeling modes
Direct controllability at intersections
Requires reliable command or route labels
Energy-based / implicit BC
Energy over observation-action pairs
Yes
Flexible multimodal scoring
Requires candidate optimization or sampling; not evaluated here
Behavior Transformer
Discrete latent/action representation
Yes
Represents several behavior modes
Quantization and architecture choices differ from continuous diffusion
Diffusion-augmented BC
Conditional BC policy plus joint state–action diffusion model
Potentially
Combines conditional and joint expert modeling
The direct policy, rather than the diffusion model, is the primary action predictor
Diffusion-BC
Conditional generative action model
Yes
Continuous stochastic action generation
Iterative inference and inconsistent closed-loop behavior in the hardest experiment
Table 1 : Qualitative positioning of the methods most directly related to this study. “Unlabeled multimodality” means that several valid actions may appear for similar observations without an explicit command identifying the desired mode
Figure 1 : Diffusion-2BC training architecture. The shared feature extractor processes the observation once. Its representation conditions the denoising branch used to compute LDBC and feeds an auxiliary deterministic head used to compute LMSE . Their weighted sum updates the shared representation. The auxiliary head is discarded after training, and test-time actions are generated only by reverse diffusion
Component
MSE-BC
Diffusion-BC
Diffusion-2BC
Visual encoder
Residual CNN
Residual CNN
Shared residual CNN
Executed head
CARLA: MLP 4608 – 2304 – 1024 – 64 – 2
Transformer denoiser
Transformer denoiser
Auxiliary head
–
–
CARLA: MLP 4608 – 2048 – 512 – 128 – 32 – 2
Inference
One deterministic pass
T reverse steps
T reverse steps
Table 2 : Compact comparison of the evaluated policy architectures
Figure 2 : Examples of the CARLA observations supplied to the policies. In (a), the yellow planned-route channel identifies the reference path, while blue and magenta encode the drivable road and lane boundaries. In (b), the route channel is omitted, so the policy must choose among locally valid road continuations. In both configurations, black denotes the non-drivable background.
Figure 3 : CARLA demonstration datasets. Each colored line is a PID-expert trajectory recorded with the corresponding observations and controls. The route-conditioned dataset supplies a visible planned-route channel. The route-free dataset omits that channel and includes multiple intersection maneuvers, so similar BEV observations can be associated with different valid actions
Experiment
Dataset
Models
Training
Evaluation
Main metrics
Generalization role
Claw environment
20,000 pairs; also 90% subset
MSE-BC, Diffusion-BC, Diffusion-2BC
Simplified supervised task
300 predictions
Mean mask-distance error
Controlled multimodality
Route-conditioned CARLA
∼ 1,900 pairs from one Town01 route
All three
Three trainings; evaluations every 50 epochs
Ten new random routes; 3,000-step limit
Distance and relative completion
Novel routes in training map
Route-free CARLA
34 Town01 routes; ∼ 16,200 pairs
All three
Two trainings; through epoch 700; evaluations every 10
Five repeated rollouts from same start
UoM distance and path diversity
Town01-selected checkpoints transferred to Town02
Table 3 : Summary of experimental protocols
Method
Full dataset
90% dataset
MSE-BC
4.02
4.25
Diffusion-BC
0.68
0.70
Diffusion-2BC, exponential α
0.63
0.69
Diffusion-2BC, fixed α=0.3
0.61
0.63
Table 4 : Results in the Claw environment. The score is the mean distance to the nearest valid mask region (lower is better), averaged over seven held-out observations with 300 sampled actions per observation
Method
Distance (m)
Relative completion
MSE-BC
1430.05±580.55
0.83±0.26
Diffusion-BC
1339.85±150.60
1.00±0.00
Diffusion-2BC
1498.09±74.81
1.00±0.00
Table 5 : Route-conditioned CARLA evaluation at the best completion-oriented checkpoints: epoch 200 for MSE-BC, epoch 300 for Diffusion-BC, and epoch 300 for Diffusion-2BC. Values aggregate three independently trained models, each evaluated on ten random routes; the reported ± values are standard deviations over the resulting 30 rollouts
Figure 4 : Route-conditioned CARLA evaluation during training. Each point combines three training runs and ten random evaluation routes per model. The curves show both progress and variability; an infraction terminates the rollout before the maximum horizon
Figure 5 : Town02 cross-map performance for five checkpoints preselected from Town01 mean-distance performance, using one Town01-trained model per architecture. Bars show mean route-free distance and bounds show the minimum and maximum across repeated rollouts. MSE-BC repeats a single deterministic path, Diffusion-BC has large variability, and Diffusion-2BC achieves the highest peak mean distance
Environment
Metric
MSE-BC
Diffusion-BC
Diffusion-2BC
Town01
Peak mean distance (UoM)
124.945
51.873
304.697
Reported min–max at peak (UoM)
identical within runs
0.907–121.469
116.741–360.201
Town02
Highest mean in transferred set (UoM)
70.563
165.717
383.963
Min–max at corresponding candidate (UoM)
identical within runs
7.489–364.232
150.107–429.379
Table 6 : Route-free multi-intersection distance results. All policies are trained in Town01, and Town02 is unseen during training. Town01 candidate checkpoints are selected by mean-distance performance; the Town02 row reports the highest Town02 mean among those five preselected candidates rather than a sweep over all training epochs.
Figure 6 : Route-free rollouts; each panel overlays ten trajectories from the same initial point. The Town01 cross-method panels use the highest mean-distance checkpoint for each architecture (epochs 60, 210, and 330 for MSE-BC, Diffusion-BC, and Diffusion-2BC). The Town02 panels show epochs 220, 230, and 680 from the fixed set of five checkpoints preselected in Town01; these are the highest observed Town02 means within that transferred set, not checkpoints obtained by a new Town02 search. The last row contains additional Diffusion-2BC checkpoints selected only for qualitative inspection of route diversity. These qualitative panels are not used to compute or select the values in Table 6 .
Supplementary Fig. S1. Reward distributions over 100 randomly generated evaluation tracks for one representative trained model of each architecture. The dashed red line marks reward 900. These distributions were previously reported in the conference article and are reproduced here to make the supplementary account self-contained.
Supplementary Fig. S2. Cumulative reward over environment steps for 20 representative evaluation episodes from one trained model of each architecture. These curves were not included in the preliminary conference article. The MSE-BC curves frequently plateau after missed track tiles, whereas most Diffusion-BC curves continue toward successful completion. The plots provide qualitative context for the aggregate statistics and are not evidence about Diffusion-2BC.
Component
Layer sequence and dimensions
Used by
Residual block 1
Two Conv2D layers, kernel 3×3 , stride 1, padding 1, 64 output channels; BatchNorm and GELU after each convolution; residual sum scaled by 1/2 ; MaxPool2D 2×2
All policies
Residual block 2
Two Conv2D layers with the same kernel, stride, padding, normalization, activation, and residual scaling; 128 output channels; MaxPool2D 2×2
All policies
Visual projection
AvgPool2D 8×8 , followed by flattening; resulting dimension depends on observation resolution
All policies
Deterministic policy head
Flattened visual vector →2304→1024→64→da with ReLU after the hidden linear layers
MSE-BC
Observation embedding
4,608-dimensional flattened visual vector →128→128 with LeakyReLU between the linear layers for CARLA
Diffusion-BC, Diffusion-2BC
Noisy-action embedding
da→128→128 with LeakyReLU between the linear layers
Diffusion-BC, Diffusion-2BC
Supplementary Table S2. Detailed CARLA neural-network components used by MSE-BC, Diffusion-BC, and Diffusion-2BC.
Parameter
Route-conditioned CARLA
Route-free CARLA
Learning rate
10−4
10−4
Learning-rate schedule
Fixed
Fixed
Feature hidden units
128
128
Batch size
32
32
Optimizer
Adam
Adam
Weight decay
0
0
Supplementary Table S3. Main training and diffusion hyperparameters for the CARLA experiments.
Supplementary Fig. S3. Town01 route-free performance for the five highest mean-distance checkpoints of one trained model per architecture. Bars show mean distance and bounds show the minimum and maximum across repeated rollouts.