Uncertainty-Gated Exploration Noise Suppresses Task Collapse in Online RL Fine-Tuning of a Flow-Matching Vision-Language-Action Policy
Organizations: Karlsruhe Institute of Technology · Çukurova University
Abstract
Online reinforcement learning fine-tuning of pretrained flow-matching vision-language-action (VLA) policies promises robots that keep learning after deployment, but continued updates often destroy competence on individual tasks while the aggregate still looks healthy. We study this failure mode, which we call task collapse, under a matched small-compute budget on LIBERO-10 with a 450M-parameter SmolVLA policy trained by PPO with stochastic (SDE) sampling. Three exploration-noise policies differ in one live variable: a fixed noise scale, a ReinFlow-style learned noise network, and an uncertainty-gated controller that redistributes exploration across task streams from task-agnostic novelty and competence signals, without task labels or episode boundaries. Under the pooled definition, fixed noise collapses tasks in two of three seeds and learned noise in every seed measured to iteration 200, while the controller collapses none in any of its three seeds. Measured parameter displacement shows the controller's action expert keeps changing, while its mean applied noise is close to the fixed scale in the available logs. The matched comparison supports the controller's effect on task preservation; the separate contributions of its adaptation across states and over time are not disentangled. A lower fixed scale slows the decline but does not stop it. No arm improves on the behavior-cloning baseline in this budget. Two properties of that regime are measured beside this result, not offered as its cause: following the reference recipe, training runs in bfloat16 with no fp32 master copy, under which 96.02% of the action expert's elements stay bit-identical across three consecutive iterations, and an fp32 master copy at the reference learning rate collapses both arms in a single-seed observation. We release tools measuring per-task collapse under four definitions, rescoring noise and instrument tares.
Figures & tables
| Arm | Seed | Last | Pool SR | Wilson 95% | Pool | Point | Sust. | Ever |
| K controller | 1337 | 200 | 0.585 | [0.555, 0.614] | 0/7 | 0/7 | 0/7 | 0/7 |
| K controller | 2337 | 200 | 0.564 | [0.534, 0.594] | 0/7 | 0/7 | 0/7 | 0/7 |
| K controller | 3337 | 200 | 0.589 | [0.559, 0.618] | 0/7 | 1/7 | 0/7 | 1/7 |
| F fixed 0.25 | 1337 | 200 | 0.316 | [0.289, 0.345] | 4/7 | 7/7 | 6/7 | 7/7 |
| F fixed 0.25 | 2337 | 200 | 0.472 | [0.442, 0.503] | 1/7 | 4/7 | 2/7 | 4/7 |
| F fixed 0.25 | 3337 | 200 | 0.510 | [0.480, 0.541] | 0/7 | 0/7 | 1/7 | 2/7 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| arm | last | pooled | last scan | sustained | ever | first crossing |
|---|---|---|---|---|---|---|
| K controller, seed 2337 | 200 | 0/7 | 0/7 | 0/7 | 0/7 | none |
| K controller, seed 3337 | 200 | 0/7 | 1/7 | 0/7 | 1/7 | 160 |
| F fixed 0.25, seed 1337 | 200 | 4/7 | 7/7 | 6/7 | 7/7 | 20 |
| F fixed 0.25, seed 2337 | 200 | 1/7 | 4/7 | 2/7 | 4/7 | 120 |
| F fixed 0.18, seed 1337 | 200 | 1/7 | 2/7 | 1/7 | 3/7 | 60 |
| F fixed 0.25, seed 3337 | 200 | 0/7 | 0/7 | 1/7 | 2/7 | 80 |
| arm | precision | last scan | pooled | Wilson 95% | first crossing |
|---|---|---|---|---|---|
| K, seed 1337 | bfloat16 in place | 200 | 614/1050 = 0.5848 | [0.5547, 0.6142] | none |
| K, seed 1337 | fp32 master copy | 200 | 297/1050 = 0.2829 | [0.2564, 0.3109] | iteration 60 |
| F 0.25, seed 1337 | bfloat16 in place | 200 | 332/1050 = 0.3162 | [0.2888, 0.3449] | iteration 20 |
| F 0.25, seed 1337 | fp32 master copy | 200 | 90/1050 = 0.0857 | [0.0703, 0.1042] | iteration 20 |
| group | elements | unchanged | unchanged % | tensors that moved |
|---|---|---|---|---|
| frozen side (VLM body) | 350165184 | 350165184 | 100.00 | 0/345 |
| action expert body | 98245840 | 94338243 | 96.02 | 112/145 |
| flow head | 1635152 | 54 | 0.00 | 10/10 |
| VLM LoRA | 819200 | 496667 | 60.63 | 124/128 |
| value head | 434944 | 1 | 0.00 | 5/5 |
| total | 451300320 | 445000149 | 98.60 |
| arm | seed | variant | window | expert | LoRA | critic | bit-ident. % |
| F | 1337 | fixed 0.25 | 100–140 | 0.0009 | 0.0048 | 0.0175 | 13.3 |
| F | 1337 | fixed 0.25 | 160–200 | 0.0009 | 0.0046 | 0.0186 | 13.3 |
| F | 1337 | fp32 master | 100–140 | 0.0034 | 0.0057 | 0.0177 | 1.4 |
| F | 1337 | fp32 master | 160–200 | 0.0034 | 0.0056 | 0.0209 | 1.4 |
| F | 1337 | fixed 0.15 | 100–140 | 0.0009 | 0.0044 | 0.0181 | 13.3 |
| F | 1337 | fixed 0.15 | 160–200 | 0.0009 | 0.0046 | 0.0156 | 13.3 |
| benchmark family | tasks | input | baselines | seeds | error bar |
|---|---|---|---|---|---|
| OpenAI Gym | 4 | state | DPPO, FQL | 3 | mean 1 SD |
| Franka Kitchen | 3 | state | DPPO, FQL | 3, 5 on two | mean 1 SD |
| Robomimic | 3 | pixel | DPPO, FQL | 3 | mean 1 SD |
| ours, LIBERO-10 | 7 | pixel and state | BC tare, three arms | 3 per arm | Wilson and run-to-run SD |
| ours, their gym | 8 of 9 | state, pixel | three arms | 3 per arm | seed SD, eval SD |
| file | line | what it sets |
|---|---|---|
| examples/embodiment/config/model/pi0.yaml | 6 | precision: null |
| examples/embodiment/config/model/pi0_5.yaml | 6 | precision: null, with the comment that it must not be changed |
| examples/embodiment/config/libero_10_ppo_openpi.yaml | 118 | precision: ${actor.model.precision} |
| examples/embodiment/config/libero_10_ppo_openpi.yaml | 152–154 | param_dtype, reduce_dtype and buffer_dtype all bound to the same variable |
| rlinf/config.py | 379–384 | use_fsdp_mixed_precision = not (all_none or all_fp32) |
| examples/embodiment/config/training_backend/fsdp.yaml | 22–31 | mixed_precision null, amp_autocast.enabled False, grad_scaler.enabled False |
| file | what it holds | how it is made |
|---|---|---|
| arm_table.md | every arm, its pooled value, both rulers, the four collapse columns | code/make_arm_table.py |
| t1_plateau_pool.md | the plateau pool of every series | code/make_figures_legacy.py |
| t5_collapse_definitions.md | the collapse definitions read side by side | code/collapse_three_ways.py |
| t6_fp32_pilot.md | the fp32 master-copy pilot | code/make_t6_reading.py |
| parameter_movement_by_arm.md | relative movement of each trainable group per arm and window | code/measure_param_movement.py |
| bf16_swallowed_updates.md | the element-level bfloat16 measurement and the fp32 arithmetic beside it | authored record, append-only journal of readings taken from checkpoint tensors |
| DPPO | WSRL | this paper | |
|---|---|---|---|
| benchmark families | 5 | 4 | 4 |
| policy classes | 1, diffusion (MLP, UNet, ViT-MLP); Gaussian and GMM as baselines | 1, soft actor-critic policy (MLP) | 2 |
| real robot | yes, Franka, 20 trials | yes, Franka peg insertion, 20 trials | no |
| named baselines | 11 | 6 in simulation, SERL on the robot | none; 3 arms against the policy’s own BC tare |
| seeds | 5 (Gym, D3IL), 3 (Robomimic, Kitchen, Furniture-Bench) | not stated | 3 per arm |
| variability reported | mean over seeds; two figures state SD not shown | no measure stated | Wilson per scoring and repeat-scoring SD |