Permutation-Robust Decision Modeling with Candidate-Independent Block-Causal Attention
Organizations: Independent Researcher
Abstract
Decision models often score a variable-sized set of candidate actions encoded in a single sequence. This setting is increasingly relevant for System 1 components inside generative systems, where candidates may be proposed or ordered differently across runs. Standard causal cross-encoding is expressive, but it can make a candidate's score depend on serialization order rather than on the underlying decision problem. We introduce candidate-independent block-causal attention, which preserves causal computation within the shared context and each candidate while blocking cross-candidate information flow and resetting candidate positions. We compare this architecture with standard causal attention and complementary invariant baselines across Gemma 3 1B, Qwen3 1.7B, and Qwen3 4B backbones. Candidate-independent attention consistently reduces permutation sensitivity while retaining competitive decision quality; ablations indicate that candidate isolation is the primary source of the effect, with position resetting completing the intended symmetry. A larger Qwen3-4B study further examines the behavior of the proposed architecture with substantially more training data. Code is available at the \href{https://github.com/guyAmit/ci-decision-models}{\textcolor{blue}{project repository}}, and the \href{https://huggingface.co/Guy-Amit/qwen3-4b-ci-decision-4096-poc}{\textcolor{blue}{Qwen3-4B model artifact}} is available on Hugging Face.
Figures & tables
| Backbone | Model | Test | OOD | |||||
|---|---|---|---|---|---|---|---|---|
| Acc. | NLL | Brier | Acc. | NLL | Brier | Flip% | ||
| Gemma 3 1B | Causal | 84.38 | .335 | .178 | 80.44 | .640 | .291 | 8.30 |
| CI | 84.66 | .327 | .171 | 80.71 | .723 | .301 | .11 | |
| Qwen3 1.7B | Causal | 81.84 | .388 | .209 | 75.09 | .706 | .358 | 18.16 |
| CI | 84.96 | .348 | .185 | 74.37 | .735 | .375 | 1.20 | |
| Qwen3 4B | Causal | 91.05 | .208 | .095 | 83.43 | .738 | .258 | 7.19 |
| Subset | Train | Val. | Val. Acc. | Val. NLL | Val. Flip |
| Open-Jev | 79,116 | 3,000 | 93.10% | 0.1721 | 0.20% |
| xLAM | 57,000 | 3,000 | 100.00% | 0.0000 | 0.00% |
| Aegis 2.0 | 30,007 | 3,000 | 97.50% | 0.0628 | 0.03% |
| ToolACE | 10,700 | 3,000 | 100.00% | 0.0002 | 0.00% |
| Repo injection | 5,387 | 3,000 | 85.03% | 0.3239 | 0.83% |
| WildGuardTrain | 67,790 | 3,000 | 98.00% | 0.0492 | 0.03% |