Behavioral Foundation Models for Quality Diversity
Organizations: Sorbonne Université, ISIR Paris, France
Abstract
Behavioral Foundation Models (BFMs) are an emerging paradigm in reinforcement learning, playing a role analogous to large language models in natural language processing: they have shown remarkable versatility, enabling zero-shot performance, fast imitation, and online adaptation, all by exploiting the structure of a latent space. In this work, we investigate whether the latent behavioral space induced by BFMs can serve as an effective search space to discover large repertoires of behaviorally diverse and high-performing policies through Quality-Diversity (QD) methods. While QD methods generally search directly in high-dimensional policy parameter space, in this paper, we present BFM-QD, a framework that performs QD search in the compact latent space of a BFM. We further show that the BFM-QD framework provides a closed-form, gradient-free policy improvement operator that approximates a policy gradient update, but requires no critic training and no backpropagation. Across continuous-control benchmarks spanning dense locomotion, sparse navigation, and contact-rich manipulation, BFM-QD consistently outperforms parameter-space baselines, with particularly stark gains in sparse and deceptive settings, where all tested parameter-space QD methods collapse to near-zero performance. These results show the effectiveness of the BFM-QD framework, benefiting from the synergy between dimensionality reduction of the search space and offline pretraining from diverse behavioral data. This positions BFMs as a general-purpose backbone for QD optimization, extending their utility beyond zero-shot task solving to the discovery of diverse behavioral repertoires.
Figures & tables
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
| Environment | Behavioral Descriptor | Fitness Function | Density | Difficulty |
| HalfCheetah | Foot contact | Run backward | Dense | Easy |
| Walk backward | Dense | Easy | ||
| Walk forward | Dense | Easy | ||
| Run forward | Dense | Easy | ||
| Walker | Foot contact | Run backward | Dense | Easy |
| Run forward | Dense | Easy |
| Parameter | Value |
|---|---|
| Archive grid size | |
| Number of generations | 500 |
| Parallel environments | 400 |
| Episode length (Brax) | 500 steps |
| Episode length (OGBench) | 1 000 steps |
| Parameter | Value |
| Mutation |
| Parameter | Value |
|---|---|
| Archive policy MLP | hidden units |
| Critic hidden width | 256 |
| Critic learning rate | |
| Actor learning rate | |
| Replay buffer size | 1 000 000 |
| Batch size | 256 |
| Parameter | Value |
|---|---|
| Offspring split (GA / PG / AI) | 50% / 25% / 25% |
| Similarity length scale | 0.1 |
| Parameter | Value |
|---|---|
| Archive policy MLP | hidden units |
| VAE latent dimension | 50 |
| VAE encoder/decoder width | 256 |
| VAE learning rate | |
| VAE training epochs | 5 |
| VAE KL weight | 0.01 |
| Parameter | Value |
|---|---|
| Offline training steps | 2 000 000 |
| Batch size | 2 048 |
| Learning rate (actor, critic) | |
| Target Polyak | 0.01 |
| Hidden width | 1024 |
| Feature dimension | 512 |
| Parameter | Value |
|---|---|
| Latent dimension | 50 |
| SF training steps | 2 000 000 |
| SF batch size | 1 024 |
| SF learning rate | |
| SF target Polyak | 0.01 |
| Actor/critic training steps | 2 000 000 |
| Parameter | Value |
|---|---|
| Latent dimension | 50 |
| Training steps | 2 000 000 |
| Batch size | 1 024 |
| Learning rate | |
| Discount | 0.99 |
| EMA | 0.01 |
| Parameter | Value |
|---|---|
| Latent dimension | 50 |
| Training steps | 2 000 000 |
| Batch size | 1 024 |
| Learning rate | |
| Target EMA | 0.01 |
| Actor std. dev. | 0.2 |
| Parameter | Value |
|---|---|
| Step size | 0.02 |
| BI probability | 0.5 |
| Gaussian mutation | 1.0 |
| z-inference batch size | 10 000 |
| Parameter | Value |
|---|---|
| Interpolation probability | 0.5 |
| distribution | |
| Gaussian mutation | 1.0 |
| Method | Walker | Cube |
|---|---|---|
| ME | ||
| DDE-Elites | ||
| ME+Pretrain | ||
| DDE-Elites+Pretrain (BC) | ||
| DDE-Elites+Pretrain (TD3) | ||
| PoMS |
| Method | Walker | Cube |
|---|---|---|
| , | ||
| , | ||
| FB-BI |
| Pretraining data | Walker | AntMaze |
|---|---|---|
| RND | ||
| QD (FB-QD) |
| Phase | Method | GPU-hours |
| BFM Pretraining (one-time) | BYOL | 2.67 |
| Lap | 2.67 | |
| TD-JEPA | 2.50 | |
| FB | 2.32 | |
| BTD-FB | 3.17 | |
| QD Search (per run) | FB-ME | 1.63 |
| Method | Backward Vel | Forward Vel | Upright Walk | Jump |
|---|---|---|---|---|
| Baselines | ||||
| MAP-Elites | 14801 514 | 15253 538 | 13746 493 | 13772 465 |
| CMA-ME | 14580 489 | 14177 507 | 13241 511 | 13262 481 |
| PGA-ME | 12410 491 | 13081 535 | 12124 474 | 12456 475 |
| DCRL-ME | 13410 413 | 14089 456 | 13714 436 | 13286 462 |
| ME+Pretrain | 13465 268 | 13576 264 | 12328 280 | 12595 268 |
| Method | Backward Run | Backward Walk | Forward Walk | Forward Run |
|---|---|---|---|---|
| Baselines | ||||
| MAP-Elites | 11331 524 | 11874 534 | 11402 490 | 11230 520 |
| CMA-ME | 12381 507 | 13200 514 | 13724 469 | 12371 534 |
| PGA-ME | 10685 463 | 14326 473 | 12523 478 | 12636 498 |
| DCRL-ME | 12466 460 | 13103 479 | 14082 456 | 12324 472 |
| ME+Pretrain | 12302 291 | 13293 338 | 13618 315 | 12425 309 |
| Method | Energy Eff. | Energy Eff. | Wrist Config | Grasp Style |
|---|---|---|---|---|
| Baselines | ||||
| MAP-Elites | 0 0 | 2 0 | 305 0 | 6 0 |
| CMA-ME | 4 0 | 2 0 | 443 0 | 29 0 |
| PGA-ME | 0 0 | 89 0 | 680 5 | 8 1 |
| DCRL-ME | 0 0 | 141 0 | 521 0 | 6 0 |
| ME+Pretrain | 3 0 | 1 0 | 152 0 | 4 0 |