cs.AISep 28, 2026

GlyphBench: A Playground for Language-Model Reinforcement Learning

Authors: Roger Creus Castanyer, Marc-Alexandre Côté, Matthew James Sargent, Augustine N. Mavor-Parker, Glen Berseth, Pablo Samuel Castro

Organizations: Mila - Quebec Artificial Intelligence Institute · Université de Montréal · Vmax

Abstract

We introduce GlyphBench, an environment suite for reinforcement learning (RL) post-training of language-model agents, with over 360 tasks spanning diverse games. GlyphBench renders spatial observations as two-dimensional Unicode grids and connects training, evaluation, and trajectory replay through a unified interface designed to support efficient and reproducible research. We use GlyphBench to study how observation interfaces, reasoning effort, and agent harnesses affect performance, and how RL configurations shape learning dynamics. Our results show that glyph observations outperform native text and pixels in our Craftax experiments, with further gains on several BALROG environments. RL on 100 GlyphBench tasks improves Qwen3.5-4B on held-out Reasoning Gym problems, reaching 63.48% accuracy and outperforming the base model, a math-trained baseline, and a code-trained baseline. These experiments provide empirical evidence that reasoning gains from gameplay can yield stronger transfer than math or code. Together, these results highlight GlyphBench's value as a testbed for systematic research on how language-model agents learn, interact, and generalize.

Figures & tables

Appendix figures & tables23 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ExTra: Exploratory Trajectory Optimization for Language Model Reinforcement Learning

    Jun 23, 2026Wenyang Hu, Junxiang Jia, Zhen Shu +3Verifiable RewardsTrajectory-Level Credit

  2. SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents

    Jun 11, 2026Ziyi Wang, Yuxuan Lu, Yimeng Zhang +8Language-Model AgentsFrontier Large Language Model Agents