We introduce GlyphBench, an environment suite for reinforcement learning (RL) post-training of language-model agents, with over 360 tasks spanning diverse games. GlyphBench renders spatial observations as two-dimensional Unicode grids and connects training, evaluation, and trajectory replay through a unified interface designed to support efficient and reproducible research. We use GlyphBench to study how observation interfaces, reasoning effort, and agent harnesses affect performance, and how RL configurations shape learning dynamics. Our results show that glyph observations outperform native text and pixels in our Craftax experiments, with further gains on several BALROG environments. RL on 100 GlyphBench tasks improves Qwen3.5-4B on held-out Reasoning Gym problems, reaching 63.48% accuracy and outperforming the base model, a math-trained baseline, and a code-trained baseline. These experiments provide empirical evidence that reasoning gains from gameplay can yield stronger transfer than math or code. Together, these results highlight GlyphBench's value as a testbed for systematic research on how language-model agents learn, interact, and generalize.
Figures & tables
Figure 1: GlyphBench connects nine game families through spatial text observations and named actions. A shared task interface supports policy evaluation, RL training, and trajectory replay. Six families form the standard track; three provide extended tasks.
Figure 2: The same Craftax state in pixels, Craftax’s original text renderer (native text), and GlyphBench’s glyph format. Native text is a comparison interface, not part of GlyphBench. We show excerpts here and both complete text observations in Figure 11 . Glyph colors aid visualization only.
Figure 3: Zero-shot performance on 303 tasks with 25 matched seeds each. Panel (a) shows task-mean returns, with interquartile ranges and median dots; (b–c) show episode length and token use over 7575 episodes per model. Tokens include prompts and completions under each model’s tokenizer.
Figure 4: Zero-shot Craftax performance across observation formats (a), harnesses (b), and reasoning efforts (c), as a percentage of maximum reward ( 100R/226 ). Each panel specifies the settings held fixed. Whiskers show 95% t intervals around five-episode means; open markers show individual episodes in (a–b). The star marks one exploratory Astra episode with Guided Memory at maximum effort.
Figure 5: Policies learn on individual tasks across all six suites. Pale traces show training reward; dark traces show centered seven-point means. Dashed lines show mean initial-policy returns. Axes vary by task. Appendix F explains the selection and shows all 25 learning curves.
Figure 6: Learning in Snake under a fixed 6,144-rollout budget. The 45 conditions vary learning rate (rows), difficulty (columns), and group size (colors). Pale traces show batch returns; solid curves show trailing 512-rollout means; dots mark final observations. Return scales match within columns. We examine higher learning rates in Appendix J .
Figure 7: One policy learns across 100 tasks over 100 gradient steps. Panel (a) shows training returns with a trailing nine-step mean; (b) evaluates the same tasks on new environment seeds.
Figure 8: Reasoning Gym accuracy at a 32K output-token budget for the base model and policies trained on games, math, or code. Each policy answers the same 3,100 problems, with three attempts per problem. Panel (a) averages accuracy across 31 tasks; (b) separates the results by category. Whiskers show 95% problem-bootstrap intervals, resampling within tasks and keeping each problem’s three attempts together.
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Suite
Tasks
Track
Task scope
Classics
50
Standard
Puzzles and familiar games for planning, memory, and control.
MiniGrid
71
Standard
Navigation, keys, doors, obstacles, and memory.
MiniHack
63
Standard
Dungeon navigation, items, combat, and skill composition.
Craftax
60
Standard
Focused survival, crafting, combat, and progression tasks.
MiniAtari
43
Standard
Arcade control with compact dynamics and short horizons.
Procgen
16
Standard
Procedural platformers, shooters, and mazes.
Appendix
Table 1: The 362 GlyphBench tasks. Standard tasks bound episodes to 512 steps and returns to [−1,1] . Extended tasks support longer episodes and are excluded from the standard aggregate score.
Figure 9: The final decision of a successful MiniAtari BankHeist episode. Replay connects the agent’s observation and reasoning to its chosen action and the resulting environment feedback.
Figure 10: The nine Craftax floors in pixels and glyphs. Each pair shows the same local state, with HUDs and legends omitted. Blank glyph regions are unobserved cells; colors aid visualization only. The paired views show how the glyph interface represents the terrain and entities encountered during progression.
Figure 11: Complete environment observations for the initial Craftax state shown in Figure 2 . (a) GlyphBench’s glyph grid, full HUD, and symbol legend. (b) Craftax’s original text renderer: all 99 cells, inventory, and status information. The map listing continues down three columns. Line wrapping and column placement are adjusted for readability.
Component
Basic
Guided Memory
Instructions
Task, core mechanics, observation conventions, and action format
Extended game guide, action format, memory protocol, and current-floor guidance
Current view
Glyph grid and standard wrapper HUD
Glyph grid and expanded inventory/status HUD
History
Growing conversation, with the last eight observation/action frames rendered in each turn
Last eight actions and rewards, plus persistent memory
External memory
No separate memory store
Six scratchpad sections and a floor-keyed landmark database
Exploration
Prior observations in the conversation
Summary of observed tiles, visited coordinates, frontiers, and remembered features
Memory calls
None
Every five steps and after relevant events; separate from action selection
Appendix
Table 2: Basic and Guided Memory in the Craftax harness comparison. Both use glyph observations, native game actions, and the selected model reasoning-effort setting. Neither requests an explicit output-token cap.
Environment
Original BALROG language interface
BabyAI
Descriptions of visible objects and their positions relative to the agent, including door state and carried objects, together with the language task.
Baba Is AI
Active rules and descriptions of objects and movable rule-word blocks using horizontal and vertical offsets from the controlled object.
Crafter
Descriptions of nearby terrain and entities by distance and direction, the object directly ahead, survival statistics, and inventory.
MiniHack
Language-wrapper descriptions of surroundings, inventory, and player statistics, supplemented by a two-dimensional ASCII map.
Appendix
Table 3: Original BALROG observation formats for the four environments in the interface comparison. These are BALROG’s environment-specific representations, distinct from GlyphBench’s shared glyph format.
Figure 12: We compare the original BALROG interfaces with glyph observations on four environments. Bars show mean progression over paired episodes; differences are percentage points. Table 3 describes each original interface, including MiniHack’s existing ASCII map.
Setting
Single-task runs
Snake
Multitask
Learning rate
10−6
Sweep
5×10−6
Action output tokens
8,192
4,096
4,096
Trainer sequence length
32,768
32,768
65,536
Trainer loss
default
ipo
ipo
Appendix
Table 4: Training settings. Output budgets apply per action-generation call; sequence length is the trainer’s token limit. Snake learning rates and group sizes are swept as described in Section 4 .
Figure 13: All 25 single-task learning curves, ordered by suite and task. Pale traces show training rewards and dark traces show centered seven-point means. Dashed lines show initial-policy rollout means where available. Endpoint ticks indicate each task’s training-step and reward range. All runs start from Qwen3.5-4B and use one training seed.
Figure 14: We evaluate the multitask policy on new environment seeds across all six training suites, using a common return scale. Each point averages task means within a suite. All suites improve over the displayed interval, while Procgen remains below zero.
Figure 15: The multitask policy learns jointly on all 100 tasks; here we show tasks 1–50, ordered by suite and task. Pale lines show task-mean batch returns; dark lines show trailing 15-gradient-step means. All panels share steps 0–250 and returns [−1,1] .
Figure 16: Training curves for the remaining 50 tasks of the same multitask policy. Axes and smoothing match Figure 15 ; together, the figures show each task once.
Figure 17: We evaluate transfer beyond gameplay by comparing performance before and after GlyphBench RL. (a) Reasoning Gym solve rates at three output budgets. (b) MathArena accuracy over all attempts. (c) SWE-bench pass rates. Bars start at zero; panel (c) spans 0–15%. Table 5 reports paired changes and uncertainty.
Benchmark
N
Base
GlyphBench RL
Δ (pp)
Paired 95% CI (pp)
MathArena aggregate
732
78.55
78.42
−0.14
[−2.46,+2.19]
AIME 2024 I
60
90.00
93.33
+3.33
[+0.00,+10.00]
AIME 2024 II
60
90.00
90.00
+0.00
[−5.00,+5.00]
AIME 2025
120
85.00
85.00
+0.00
[−5.00,+5.00]
AIME 2026
120
85.00
85.00
+0.00
[−5.83,+5.83]
HMMT Feb 2025
120
70.00
64.17
−5.83
[−14.17,+2.50]
Appendix
Table 5: External evaluation of the base and GlyphBench RL models. Rates are percentages; changes and paired 95% intervals are percentage points. N counts attempts for mathematics and reasoning, and instances for coding. Resampling clusters attempts by problem.
Figure 18: Reasoning Gym performance before and after GlyphBench RL by category and output budget, on a common 0–100% scale. Every category improves at 8K and 16K; the smaller 2K budget yields a weaker and less consistent benefit. Table 6 reports the corresponding values.
Figure 19: Changes after GlyphBench RL on external benchmarks, in percentage points. The left panel shows aggregate changes with paired 95% bootstrap intervals; the right separates Reasoning Gym changes by category and output budget. Positive estimates whose intervals include zero do not establish an improvement.
Category
2K: base / ours ( Δ )
8K: base / ours ( Δ )
16K: base / ours ( Δ )
Algebra
40.61 / 46.78 ( +6.17 )
64.72 / 77.06 ( +12.33 )
71.00 / 79.22 ( +8.22 )
ARC
5.22 / 4.33 ( −0.89 )
19.22 / 34.78 ( +15.56 )
38.22 / 52.89 ( +14.67 )
Cognition
24.29 / 25.62 ( +1.33 )
36.05 / 38.00 ( +1.95 )
37.86 / 40.57 ( +2.71 )
Games
3.38 / 4.92 ( +1.54 )
31.42 / 39.58 ( +8.17 )
47.88 / 60.08 ( +12.21 )
Geometry
19.33 / 19.00 ( −0.33 )
48.50 / 51.50 ( +3.00 )
50.67 / 53.33 ( +2.67 )
Graphs
32.53 / 35.00 ( +2.47 )
64.13 / 80.20 ( +16.07 )
78.20 / 86.93 ( +8.73 )
Appendix
Table 6: Reasoning Gym solve rates (%): base / GlyphBench RL, with changes in percentage points. Categories contain 6, 3, 7, 8, 2, and 5 tasks in row order, each with 100 problems and three attempts per problem.
Figure 20: Math and code RL improve rewards in their respective training domains. The top panels show training reward (pale), trailing nine-step means (solid), and held-out domain reward (dashed). The bottom panels show average output tokens on held-out problems from each domain. These rewards measure learning within each domain; we evaluate reasoning transfer separately on Reasoning Gym.
Reference policy
Accuracy difference
Paired 95% interval
Base
+7.44
[+6.55,+8.34]
Math RL
+1.19
[+0.38,+2.02]
Code RL
+4.51
[+3.62,+5.40]
Appendix
Table 7: GlyphBench RL minus each comparison policy at the 32K output budget. Differences and paired 95% bootstrap intervals are percentage points.
Figure 21: We compare Reasoning Gym accuracy at additional output-token budgets. Each cap uses independently sampled responses to the same problems. Whiskers show 95% problem-bootstrap intervals; all three attempts for a problem remain in the same resampled cluster.
Figure 22: Reasoning Gym generation diagnostics. (a) Problems with at least one successful completion among three attempts. (b) Responses ending at the output-token limit. (c) Mean output tokens per completion. Truncated responses remain in the denominator and can still count as correct if the answer is extractable.
Figure 23: Peak versus final observed Snake return at 10−6 , 10−5 , and 10−4 within 6,144 rollouts. Both coordinates use trailing 512-rollout means. Colors indicate group sizes, shapes indicate learning rates, and gray paths connect increasing group sizes within a rate. The diagonal marks full retention of the peak; vertical stems show the drop to final return. Gold stars mark each difficulty’s largest peak among plotted conditions. The 10−4 sweep is exploratory.
Figure 24: Snake training over 25,600 rollouts at 10−6 , comparing groups 32 and 512 across three difficulties. Pale traces show batch returns, solid curves show trailing 512-rollout means, and dots mark final observations. The dotted line marks 6,144 rollouts for reference. Group 32 achieves the higher final return on all three tasks.
Reinforcement Learning with Verifiable Rewards (RLVR) for language-model reasoning can fail at both extremes of task difficulty: easy prompts often produce all-correct, low-diversity rollout groups with little gradient signal, while hard prompts can produce all-incorrect groups with no positive reward. We introduce ExTra (Exploratory Trajectory Optimization), a GRPO-compatible framework that extracts exploration signals from the model's own rollouts. ExTra combines two mechanisms: (i) a novelty reward that adds embedding-based diversity bonuses after GRPO normalization, rewarding diverse correct solutions; and (ii) entropy-guided prefix regeneration, which scores partial trajectories using entropy signals and continues exploration from promising intermediate steps. Across six mathematical reasoning benchmarks, ExTra improves Qwen3-1.7B over GRPO by about +5 points on pass@1 and +7 points on pass@16, showing that trajectory-level exploration signals can improve both single-sample accuracy and inference-time coverage.
Language model agents are increasingly effective in solving realistic tasks through multi-turn tool use. However, training reliable tool-using agents remains challenging in practice. While reinforcement learning provides an on-policy paradigm for improving agents from their own environment interactions, its effectiveness depends heavily on the training task distribution. When tasks are fixed before training, the task distribution can become increasingly mismatched with the policy's evolving capabilities, causing many rollouts to be spent on uninformative tasks. We propose SENTINEL, a failure-driven reinforcement learning framework that turns the Solver's rollout failures into targeted training tasks. SENTINEL follows a Controller--Proposer--Solver loop: the Controller analyzes failed trajectories and summarizes recurring error patterns, the Proposer generates executable tasks that stress these weaknesses, and the Solver is trained on the targeted tasks. On Tau2-Bench Retail with Qwen3-4B-Thinking-2507, SENTINEL improves Pass^{}1 from 66.4 to 74.9 and outperforms RL on general synthetic tasks across Pass^{}k metrics. These results demonstrate that model failures provide an effective and scalable source of targeted training signal for improving tool-using language model agents.
Ziyi Wang, Yuxuan Lu, Yimeng Zhang +8
Northeastern University · Independent Researcher · Northwestern University
Reasoning Language Models (RLMs) achieve their strongest performance when they reason in English, the language for which reasoning-oriented training data is most abundant. However, reasoning trace is a clue for model interpretability and safety, and useful in practice for both the model users and for model developers. Thus, it is desirable to be able to develop a model that reasons in a language of the user's choice, while still maintaining strong reasoning performance. To this end, we study the feasibility of training a model that reasons in Japanese. We develop a Japanese-reasoning variant of Qwen-3-Swallow-8B, which is a Japanese LLM continually pretrained from Qwen-3-8B, with GRPO and evaluate it across coding, math, and science benchmarks. The study shows that reasoning-language control is feasible by training a Japanese continually pretrained model with GRPO. However, its performance is at best on par with strong English-reasoning baselines on several benchmarks. We also evaluate the trained model on Japanese cultural benchmarks and observe that the model's performance is worse than the baseline models, suggesting that the reasoning in Japanese does not immediately improve performance on culturally relevant tasks for free.