Game-Guided Skill Discovery through Self-Play for Playable Agent Control
Organizations: Georgia Institute of Technology · Simon Fraser University and NVIDIA
Abstract
We present Game-Guided Skill Discovery (GGSD), a framework that uses self-play in games to discover motor skills that are directly playable by humans. Playable skills provide a compact abstraction for controlling embodied agents through a small set of learned behaviors rather than low-level actions. To be effective, these skills should be semantically distinct, interpretable, and expressive; properties that existing unsupervised skill-discovery methods often fail to achieve simultaneously. GGSD achieves these desiderata by grounding skill discovery in competitive gameplay. A hierarchical agent competes against its past selves, with a high-level policy selecting from a small discrete skill set and a skill-conditioned low-level policy learning the corresponding behaviors. After training, a human can replace the high-level policy and directly control the agent through the same discrete skills. Despite the small number of high-level actions, skill transitions give rise to emergent combo behaviors, expanding expressivity beyond individual primitives. Across Ant, Franka-arm, and Unitree G1 environments, we show that GGSD produces human-playable skills that humans can compose to solve unseen tasks, such as Maze and CubePush, without additional training. An interactive demo is available at https://ggsd-demo.github.io.
Figures & tables
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Environment | # Skills | ||
|---|---|---|---|
| Ant Sumo | 35 | 91 | 5 |
| Ant Fencing | 35 | 91 | 5 |
| Franka Hockey | 23 | 61 | 5 |
| G1 Boxing | 86 | 194 | 6 |
| Group | Component | Dim. | |
| Me | Base height | 1 | ✓ |
| Base linear velocity (body frame) | 3 | ✓ | |
| Base angular velocity (body frame) | 3 | ✓ | |
| Base yaw | 1 | ✓ | |
| Projected gravity (body frame) | 3 | ✓ | |
| Joint positions (normalized) | 8 | ✓ |
| Term | Description | Expression | Weight |
| Approach | Decrease in distance to the opponent | 2.0 | |
| Push | Decrease in the opponent’s distance to the nearest edge | 2.0 | |
| Win | Opponent leaves the arena or falls (once, terminal) | ||
| Lose | Opponent wins | ||
| Timeout | Draw at 10 s (once, terminal) | ||
| Alive, escape, upright, action, joint-velocity | 0 | ||
| Term | Description | Expression | Weight |
|---|---|---|---|
| Approach | Decrease in distance to the opponent | 2.0 | |
| Win | Front foot hits the opponent’s torso ( N), or opponent leaves the arena / falls | ||
| Lose | Opponent’s front foot hits the agent’s torso, or agent leaves the arena / falls | ||
| Timeout | Draw at 10 s |
| Group | Component | Dim. | |
| Me | Arm joint positions | 7 | ✓ |
| Arm joint velocities ( ) | 7 | ✓ | |
| EE position (table frame) | 3 | ✓ | |
| EE linear velocity (table frame) | 3 | ✓ | |
| EE yaw | 2 | ✓ | |
| EE yaw rate | 1 | ✓ |
| Term | Description | Expression | Weight |
| Puck progress | Forward puck displacement toward the opponent’s goal (backward motion ignored) | 10.0 | |
| Goal | Agent scores (once per goal) | ||
| Concede | Opponent scores (once per goal) | ||
| Match win | Two goals first, or leading at timeout / idle-puck termination | ||
| Match lose | Mirror of match win | ||
| Draw | Tied at termination | 0 |
| Group | Component | Dim. | |
| Me | Base yaw | 1 | ✓ |
| Base height | 1 | ✓ | |
| Base linear velocity (body frame) | 3 | ✓ | |
| Base angular velocity (body frame) | 3 | ✓ | |
| Projected gravity (body frame) | 3 | ✓ | |
| Glove positions (heading frame, ) | 6 | ✓ |
| Term | Description | Expression | Weight |
| Punch | HP drained from the opponent minus HP lost | 125 | |
| Facing | Torso and pelvis facing the opponent, in | 0.03 | |
| Facing velocity | Penalises moving sideways relative to the body’s facing | 0.15 | |
| Approach | Decrease in distance to the opponent; zero when m | 3.0 | |
| Joint velocity | Squared joint velocity, clipped at rad/s | ||
| Joint limit | Excess beyond 80 % of the soft joint range, clipped at 0.3 | 0.5 |
| Hyper-parameter | Ant Sumo | Ant Fencing | Franka Hockey | G1 Boxing |
| PPO | ||||
| Parallel environments | 4096 | |||
| Rollout length (steps / env) | 32 | 32 | 24 | 32 |
| Training iterations | 70 000 | |||
| Learning epochs / mini-batches | 5 / 4 | |||
| Learning rate | ||||