GrammarRL: Effective Grammar-Constrained Decoding via Reinforcement Learning
Organizations: University of Catania · ISTC - National Research Council, Italy · University of Bologna
Abstract
Grammar-constrained generation guarantees syntactic validity, but can substantially degrade semantic quality when the model's preferred outputs are poorly aligned with the imposed grammar. This trade-off is particularly severe when the prompt is underspecified or the model has limited instruction-following ability. Beam search can partially mitigate these failures by exploring multiple valid sequences, but its computational cost grows with beam width, while sequence-level probability is only an imperfect proxy for semantic quality. We introduce GrammarRL, a label-free reinforcement learning method that adapts language models to grammar constraints without requiring annotated data. GrammarRL optimizes the model using two complementary self-supervised rewards derived from its own likelihoods: a direct reward, measuring how likely the constrained output is given the input, and a reverse reward, measuring how well the input can be reconstructed from the generated output. We optimize these rewards with a Reinforce Leave-One-Out (RLOO) objective over groups of grammar-constrained rollouts, augmented with the top-1 beam-search hypothesis and regularized towards a frozen base model. We evaluate GrammarRL on sign language gloss translation, hierarchical text classification, and named entity recognition using Llama models ranging from 1B to 8B parameters. GrammarRL consistently outperforms constrained greedy decoding, with an average improvement of 9.8 points and gains of up to 22.8 BLEU. It matches or outperforms beam search on two of the three tasks while preserving greedy-decoding inference cost. Ablations further show that the two rewards are complementary: either reward alone can underperform the untrained baseline, whereas their combination consistently improves upon it.
Figures & tables
| Model | Config | Gloss (BLEU) | WoS (Hier. F1) | CoNLL (F1-micro) |
|---|---|---|---|---|
| 1B | baseline (greedy) | 0.342 0.005 | 0.384 0.006 | 0.435 0.009 |
| baseline (beam) | 0.390 0.018 | 0.385 0.007 | 0.590 0.006 | |
| GrammarRL (reverse only) | 0.554 0.009 | 0.309 0.006 | 0.548 0.005 | |
| GrammarRL (direct only) | 0.456 0.005 | 0.288 0.007 | 0.597 0.010 | |
| GrammarRL | 0.570 0.014 | 0.394 0.008 | 0.578 0.006 | |
| GrammarRL (beam) | 0.541 0.013 | 0.391 0.008 | 0.598 0.006 |
| Method | Cost | Gloss (BLEU) | WoS (Hier. F1) | CoNLL (F1-micro) |
|---|---|---|---|---|
| Constrained greedy | 0.339 0.002 | 0.388 0.000 | 0.434 0.006 | |
| Constrained beam ( ) | 0.380 0.022 | 0.386 0.002 | 0.586 0.009 | |
| Sample-Rerank ( ) | 0.346 0.000 | 0.430 0.009 | 0.488 0.007 | |
| Full IS ( ) | 0.302 0.002 | 0.442 0.002 | 0.511 0.010 | |
| Full SMC ( ) | 0.236 0.004 | n/a | 0.578 0.009 | |
| GrammarRL ( ) | (+ training) | 0.554 0.008 | 0.400 0.010 | 0.577 0.009 |
| Lambda value | Gloss (BLEU) | WoS (Hier. F1) | CoNLL (F1-micro) |
|---|---|---|---|
| 0 (direct only) | 0.456 0.005 | 0.288 0.007 | 0.597 0.010 |
| 0.25 | 0.541 0.004 | 0.407 0.009 | 0.625 0.008 |
| 0.5 | 0.570 0.014 | 0.394 0.008 | 0.578 0.006 |
| 0.75 | 0.576 0.013 | 0.367 0.009 | 0.572 0.005 |
| 1 (reverse only) | 0.554 0.009 | 0.309 0.006 | 0.548 0.006 |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| 1B | 3B | 8B | |||||
|---|---|---|---|---|---|---|---|
| Task | |||||||
| Gloss | 1.1378 | 0.9731 | 0.6354 | 0.5493 | 0.4564 | 0.4442 | |
| 1.0693 | 1.0001 | – | – | – | – | ||
| 1.1245 | 1.0229 | – | – | – | – | ||
| 1.1147 | 0.9125 | – | – | – | – | ||
| 1.1454 | 0.9453 | – | – | – | – | ||
| Gloss | WoS | CoNLL | |
| Data | |||
| Training examples | 2k | 4k | 4k |
| Test sets | |||
| Few-shot pairs / hints | 20 / 50 | 1 / — | 10 / 80 |
| Optimization | |||
| Prompts per step (batch) | 4 | 8 | 8 |
| Task | Few-shot pairs | Dynamic hints per query | Static system content |
|---|---|---|---|
| Gloss | 20 | 50 gloss terms | — |
| WoS | 1 | — | label hierarchy (7 parents, 134 children) |
| CoNLL | 10 | entity spans | — |
| Task | Instruction | Context | Output format |
| Forward system prompt | |||
| Gloss | Simplify the sentence into basic components. Separate components with a single space; preserve the meaning and structure of the original. | 50 similar gloss terms (retrieved) | TERM TERM TERM … upper-case content words, space-separated |
| WoS | Classify the abstract into one parent–child path of the hierarchy. Use labels from the hierarchy only. | label hierarchy (7 parents, 134 children; static) | {"parent": ..., "child": ...} |
| CoNLL | Extract named entities from the sentence. Copy the first span of each type verbatim; empty string if absent; emit the JSON object only. | 4 20 similar entity spans (one line per type, retrieved) | {"person": ..., "organization": ..., "location": ..., "misc": ...} |
| Reverse system prompt (replaces the forward instruction; no context block) | |||
| Gloss | Given the glossed form, reconstruct the original English sentence as it was actually written. Output only that sentence. | ||
| Gloss (BLEU) | WoS (Hier. F1) | CoNLL (F1-micro) | |
| Sampled sequences ( ) | |||
| 3 | 0.570 0.014 | 0.394 0.008 | 0.578 0.006 |
| 6 | 0.584 0.017 | 0.403 0.008 | 0.577 0.007 |
| 9 | 0.589 0.010 | 0.402 0.009 | 0.572 0.007 |
| Beam width ( ) | |||
| 3 | 0.570 0.014 | 0.394 0.008 | 0.578 0.006 |
| THEREFORE-DESC | OFTEN-DESC | THERE-DESC | bare X | |
|---|---|---|---|---|
| 0 (direct only) | 7 | 0 | 33 | 115 |
| 0.25 | 13 | 0 | 46 | 84 |
| 0.5 | 18 | 1 | 67 | 79 |
| 0.75 | 24 | 1 | 90 | 54 |
| 1 (reverse only) | 36 | 1 | 140 | 51 |
| Task | Rule | Violation | Penalty | Check on training gold (500 rows) |
|---|---|---|---|---|
| CoNLL | Grounding | An entity string is not a verbatim substring of the sentence. | each | fires on 0 |
| CoNLL | Word boundary | The string occurs in the sentence but only inside longer words. | each | fires on 0 |
| CoNLL | Duplicate | The same non-empty string is emitted in two different fields. | per extra field | fires on 0 |
| Gloss | Precision | A completed gloss term is not grounded: after removing the -DESC / -X marker, its lemma matches no source word (suffix stripping and a dictionary lemmatizer) and it is not a closed-class function term (e.g. be, have, not, the ). | each, at most 5 | fires on 14.6% of rows |
| Gloss | Coverage | A source content word (alphabetic, at least three letters, not a stop word) is not matched by any gloss term. Judged on the complete output only. | each, at most 5 | gold covers of content words (shuffled control: ) |
| Gloss | Degenerate run | A maximal run of at least three identical adjacent terms (a run of two, as in the gold BE BE , is legitimate). | per run | fires on 0 |
| Method | What it does | Search at inference | Efficient potential | Expensive potential | Output selection | Cost |
|---|---|---|---|---|---|---|
| Constrained greedy | Mode of the locally renormalized masked policy equation 1 : one valid token at a time; no look-ahead. | None (single pass). | Grammar as a token mask (Grammar-LLM). | None. | Arg-max token at each step. | |
| Constrained beam ( ) | Searches for a higher-probability sequence under ; a mode-seeking approximation of . | Beam search, 3 hypotheses. | Same mask (Grammar-LLM). | None. | Top-scoring finished beam. | |
| GrammarRL (ours) | Amortizes sequence-level preferences into the model: LoRA policy trained with RLOO on the label-free bidirectional reward, then decoded greedily. | None at test time (greedy). | Same mask, active in training and testing. | Used only in training: direct reverse reward from the frozen model equation 6 . | Arg-max token at each step. | (+ one-off training) |
| Sample-Rerank ( ) | Best-of- under an external score: i.i.d. grammar-constrained samples, keep the one with the highest reverse likelihood of the input. No weight correction, no resampling. | independent samples. | Grammar mask (proposal). | (noisy-channel), at completion. | of the reverse score among complete samples. | generation scoring passes |
| Full IS ( ) | Importance sampling from the grammar-constrained LM towards ; weights correct the myopia of local masking; no resampling. | independent weighted samples. | Grammar mask inside AWRS, with importance weight. | (noisy-channel), applied once at completion. | Posterior mode: text with the largest total normalized weight. | Same as Sample-Rerank |
| Full SMC ( ) | Sequential Monte Carlo ( Loula et al., 2025 ) : particles extended in parallel, resampled when the effective sample size drops, steered at every step by a task-specific expensive potential. | particles, resampling if . | Grammar mask inside AWRS, with importance weight. | (CoNLL, Gloss): grounding, coverage, boundary, repetition; on prefixes and at completion. | Posterior mode over particles. | generation; potential evaluated on CPU |