Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL
Organizations: Patronus AI
Abstract
Recent growth in reinforcement learning (RL) has surfaced a need for diverse, specialized training environments. Hand-curated environments with fixed task and reward difficulties become ineffective signals as model performance improves, and sparse rewards over long horizons induce mode collapse on specific workflows or tool structures. World models that simulate environment states have matched pure rollout performance, making them promising for scaling diversity on-demand. However, autoregressive (AR) world models suffer from a left-to-right bias preventing conditioning on globally interdependent state anchors such as tool schemas, prior turns, and expected outcomes. We (i) formalize text-based world modeling as a steerable transition-dynamics problem decomposed into initial state, task context, tool schemas, domain rules, and steering directives, and (ii) curate 239,403 grounded state-action trajectories spanning nine open-source environments and twelve frontier model families. We compare AR LMs and masked diffusion language models (MDLMs), showing MDLMs, via bidirectional anchor-aware denoising, achieve better coherence, groundedness, and empirically validated rollout diversity than LLMs over 4x their parameter size, at comparable inference latency. We introduce a plug-and-play GRPO training framework with deterministic state checks, and perform zero-shot transfer ablations on three OOD environments (ScienceWorld, ALFWorld, AppWorld) across three 1.2B-7B agent backbones (LFM2.5, Qwen3, Mistral), achieving up to 47% absolute gains over baselines without environment-specific fine-tuning. We further conduct behavioral analysis of failure modes under adversarial scenarios and human evaluation on realism, outcome correctness, and training utility. We open-source our work to encourage research in this direction.
Figures & tables
| In-domain | API-Bank | OccuBench | Intercode (SQL) | |||||||||
| Model | B-1 | R-L | MAUVE | B-1 | R-L | MAUVE | B-1 | R-L | MAUVE | B-1 | R-L | MAUVE |
| Baseline Models | ||||||||||||
| Qwen-3-8B | .136 | .286 | .148 | .453 | .421 | .301 | .524 | .523 | .705 | .406 | .397 | .699 |
| Qwen-3.5-27B | .152 | .221 | .187 | .536 | .611 | .505 | .746 | .755 | .810 | .516 | .614 | .799 |
| Qwen-3.5-35B-A3B | .125 | .233 | .189 | .488 | .585 | .499 | .667 | .680 | .811 | .522 | .598 | .780 |
| Nemotron-3-Nano-30B | .120 | .152 | .223 | .412 | .316 | .510 | .675 | .667 | .822 | .486 | .465 | .785 |
| Model | Self-BLEU ( ) | Distinct-N ( ) | MAUVE ( ) |
|---|---|---|---|
| Qwen3.5-35B-A3B | .690 | .253 | .889 |
| Qwen3.5-27B | .670 | .321 | .864 |
| GPT-OSS-20B | .659 | .228 | .867 |
| SDAR-8B | .601 | .385 | .902 |
| Metric | Mean | Krippendorff’s |
|---|---|---|
| Realism | 4.75 | .932 |
| Outcome Correctness | 4.25 | .891 |
| Training Utility | 4.50 | .901 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Score | Label | Description |
| Training utility | ||
| 5 | Excellent | Correct and consistent outcome, realistic values, i.e. ideal training signal |
| 4 | Good | Minor value errors, but correct outcome type |
| 3 | Mediocre | Correct outcome type, but wrong schema or implausible values |
| 2 | Poor | Wrong outcome type or severely hallucinated structure |
| 1 | Useless | Completely wrong, nonsensical, or would teach incorrect behaviours, including incomplete outputs |
| Configuration | BLEU-1 | ROUGE-L | MAUVE | Latency ( ) |
|---|---|---|---|---|
| Qwen3.5-35B-A3B (no thinking) | .621 | .664 | .899 | 211s ( ) |
| + Thinking ( ) | .713 | .768 | .940 | 793s ( ) |
| + Thinking ( ) | .710 | .758 | .934 | 1028s ( ) |
| SDAR-8B (no thinking) | .749 | .740 | .902 | 380s ( ) |
| WeDLM-8B (no thinking) | .729 | .761 | .910 | 145s ( ) |
| Failure Mode | Env | Trajectory Snippet |
|---|---|---|
| (i) Incorrect pagination | AW | Agent requests page 1 of recommendations; WM returns the same records as page 0. Downstream, the agent double-counts songs and computes incorrect aggregates. T4: show_recommendations({page_index:0}) {total:162, songs:[{id:40},{id:41},…]} T5: show_recommendations({page_index:1}) {total:162, songs:[{id:40},{id:41},…]} same IDs |
| (ii) Object hallucination | SW | WM invents code-execution errors and objects absent from the environment state. Agent receives fabricated feedback that does not correspond to any valid ScienceWorld response. Agent: open cupboard WM: ‘‘Error during execution: name ‘cupboard_is_opened’ is not defined’’ fabricated [2pt] Agent: activate solar panel WM: ‘‘Solar panels cannot be used directly.’’ not a ScienceWorld response |
| (iii) JSON-wrapped responses | SW | WM outputs structured JSON instead of plain-text environment observations. Downstream parser receives role/content wrappers that corrupt the conversation history. Agent: look around WM: [{‘‘role’’:‘‘tool’’, ‘‘content’’:‘‘This room is called the kitchen. In it, you see: the agent, a substance called air, a chair…’’}] JSON wrapper |
| (iv) Dropped instance numbers | ALF | WM substitutes generic class names for specific receptacle instances, breaking downstream state tracking and action validation. Agent: go to toaster 1 WM: ‘‘You arrive at toaster 1. On the table , you see a bread 2.’’ Expected: ‘‘On the toaster 1 , you see a bread 2.’’ [2pt] Agent: go to microwave 1 WM: ‘‘You arrive at microwave 1. On the table , nothing…’’ Expected: ‘‘On the microwave 1 , you see nothing.’’ |
| (v) State inconsistency | SW | WM confirms picking up an object that is already in inventory, creating phantom duplicates. Agent “picks up” the thermometer three times across 17 turns. T6: pick up thermometer ‘‘You move the thermometer to the inventory.’’ T15: pick up thermometer ‘‘You move the thermometer to the inventory.’’ already held T17: pick up thermometer ‘‘You move the thermometer to the inventory.’’ third time |
| (vi) Verbose / garbled output | ALF, SW | WM generates excessively long descriptions with hallucinated items, or garbles object names by dropping instance numbers from object listings. ALFWorld: Agent: go to fridge 1 WM: ‘‘…On the fridge, you see a 2 , and a mug.’’ (object class dropped) [2pt] ScienceWorld (Qwen WM): Agent: look around (greenhouse) WM: ‘‘…a decorative stone gazing ball resting atop one corner… a rustic lantern holding dried lavender bundles… finally tucked away quietly within deepest recesses thereof lies forgotten treasure trove comprising ancient maps …’’ (886 hallucinated tokens) |
| In-domain | API-Bank | OccuBench | Intercode (SQL) | |||||||||
| Model | B-1 | R-L | MAUVE | B-1 | R-L | MAUVE | B-1 | R-L | MAUVE | B-1 | R-L | MAUVE |
| 3-shot Baseline Models | ||||||||||||
| Qwen-3-8B | .379 | .506 | .348 | .578 | .509 | .365 | .559 | .544 | .788 | .532 | .523 | .833 |
| Qwen-3.5-27B | .509 | .578 | .588 | .611 | .660 | .600 | .731 | .772 | .834 | .585 | .623 | .849 |
| Qwen-3.5-35B-A3B | .334 | .507 | .590 | .621 | .663 | .634 | .744 | .781 | .832 | .589 | .662 | .851 |
| GLM-4.7-Flash | .401 | .565 | .629 | .686 | .689 | .707 | .746 | .742 | .838 | .612 | .597 | .822 |
| Model | Architecture | B-1 | R-L | MAUVE |
|---|---|---|---|---|
| RWKV-7B-g1 | RWKV (linear attention) | .334 | .496 | .651 |
| Falcon-H1R-7B | Mamba-attention hybrid | .541 | .661 | .807 |
| Qwen-3-8B | Transformer (AR) | .636 | .682 | .873 |
| SDAR-8B | Transformer (MDLM) | .749 | .740 | .902 |
| Qwen3.5-27B (AR) | SDAR-8B (forced) | SDAR-8B (self-committed) | |
| 1 | .86 | .90 | .89 |
| 2 | .86 | .91 | .90 |
| 8 | .86 | .93 | .89 |
| 16 | .86 | .94 | .91 |
| Failure mode | WM prediction | Expected |
|---|---|---|
| (vii) API-key corruption | GetUserToken "p9o8i7u6y5t…tttepmw" | "p9o8i7u6y5t4r3e2" |
| (viii) Type coercion | SetVolume {"level": "7"} | {"level": 7} (schema type int ) |
| (ix) Web-search fabrication | “Team SoloMid won the 2013 NA LCS Summer Finals …” | “Cloud9 won the 2013 NA LCS Summer playoffs …” |
| (x) Error cascade | After one earlier error, every retry of move_to_node(…) TimeoutError | First attempt succeeds |