Long-horizon reasoning demands access to earlier information at a manageable generation cost. Full-history attention incurs growing storage and computation, while recurrent compression can lose precise details. Therefore, we present YANchor-4B, a general-purpose recurrent model that preserves crucial memory as ANchors for retrieval during subsequent reasoning. Beyond O(N)-time generation and O(1) memory, YANchor enables effective long-horizon reasoning through its multidimensional memory mechanism. For example, on challenging math problems, it achieves 82.93% mean pass@1 on AIME 2024--2026 and 63.64% on HMMT, substantially outperforming linear-time, constant-state counterparts, including larger models. It also delivers several-fold higher batched long-generation throughput than Transformer and hybrid baselines on H100. Furthermore, evaluations across dozens of benchmarks demonstrate YANchor's superiority in general-purpose capabilities.
Figures & tables
Figure 1: YANchor-4B and 11 bounded-state baselines on mathematics, code, and science. Models are grouped by parameter scale. Avg. denotes an equal-weight benchmark mean: AIME covers 2024–2026; Code covers HumanEval, MBPP, and LiveCodeBench v6; Science covers GPQA-Diamond, ARC-C, and ARC-E. HMMT uses February 2026. Section 5 defines the primary metrics.
Figure 2: Single-H100 inference: sustained decoding speed, higher batched throughput, and bounded memory demand. Left: cumulative decode speed. Middle: end-to-end output throughput, selecting the batch size with the highest measured throughput for each model and input length. Right: batch-512 peak memory demand; solid curves use measured components, dashed curves project KV growth, and the dotted line marks device capacity. Appendix D gives the workloads and runtime settings.
Figure 3: Architecture of YANchor-4B. Recurrent layers are interleaved with local-attention and memory modules. The inset shows how explicit records are written, retained as memory anchors, and retrieved through an independent read path. All history paths have fixed capacity. Image credit: NASA.
Figure 4: Causal memory maintenance. An independent writer produces K/V representations, and a causal network supplies group-specific admission scores. Each 256-token block selects a persistent Top-1,024 set per KV group. Retained records preserve their written representations as memory anchors for subsequent retrieval.
Quantity
Value
Text-model parameters
4.58B
Independent memory parameters
372M
Backbone layers
24 GDN + 8 SWA
Residual width
2,560
Memory modules
8
Memory query heads / KV groups / head dimension
16 / 4 / 256
Table 1: Architecture of YANchor-4B. Parameter totals describe the text model; the visual encoder is separate.
History mechanism
Total decode work
Persistent context state
Full-history attention
O(N2)
O(N)
Fixed-dimensional recurrence
O(N)
O(1)
Fixed-window attention
O(N)
O(1)
YANchor recurrence + SWA + memory
O(N)
O(1)
Table 2: Autoregressive scaling with generated length N .
Figure 5: The M0, R, S, and RLVR training sequence for YANchor-4B.
Model
GSM8K [ 25 ]
MATH [ 26 ]
AIME ’24 [ 27 ]
AIME ’25 [ 27 ]
AIME ’26 [ 27 ]
HMMT ’26 [ 28 ]
YANchor-4B
94.45
97.49
86.72
77.45
84.64
63.64
ARWKV-R1 1.5B [ 16 ]
21.00
12.60
0.00
0.00
0.00
1.89
RWKV-7 G1i 1.5B [ 14 ]
60.00
28.20
0.00
0.42
0.42
0.00
RWKV-7 G1j 1.5B [ 14 ]
57.60
31.40
1.25
2.08
0.00
0.00
RWKV-7 G1i 2.9B [ 14 ]
71.30
43.80
1.25
1.67
2.08
0.76
RWKV-7 G1j 2.9B [ 14 ]
73.70
48.20
0.42
4.58
1.67
0.38
Table 3: Mathematics among linear-time, constant-state models. MATH denotes MATH-500 and HMMT denotes February 2026. Scores are percentages.
Figure 6: Competition performance and generation behavior. Left: benchmark score against mean output length; each line connects YANchor and Qwen3.5-4B on the same benchmark. Right: YANchor’s normal-termination rate across the four competition sets. Accuracy measures successful solutions; normal termination measures reaching EOS before a stopping limit.
Benchmark
YANchor-4B
Qwen3.5-4B [ 1 ]
AIME 2024 [ 27 ]
86.72
89.58
AIME 2025 [ 27 ]
77.45
80.83
AIME 2026 [ 27 ]
84.64
87.92
HMMT Feb 2026 [ 28 ]
63.64
65.91
MATH-500 [ 26 ]
97.49
98.00
AIME three-year mean
82.93
86.11
Table 4: Competition mathematics and MATH-500. YANchor scores are mean pass@1 in percent; the AIME mean weights the three editions equally.
Figure 7: Cross-architecture comparison of YANchor-4B with the original Qwen3.5-4B, a Gated DeltaNet/global-attention hybrid, across mathematics, knowledge, code, and instruction following. Scores use the primary metrics defined in Section 5 .
Model
AIME ’24 [ 27 ]
AIME ’25 [ 27 ]
AIME ’26 [ 27 ]
MMLU-Pro [ 30 ]
LCB v6 [ 43 ]
IFEval [ 44 ]
YANchor-4B
86.72
77.45
84.64
79.60
61.63
88.03
Qwen3.5-4B [ 1 ]
89.58
80.83
87.92
75.90
76.20
90.02
Gemma4-E2B-it [ 5 ]
37.50
32.50
33.33
38.70
60.40
81.89
MiniCPM5-2B [ 3 ]
87.08
85.42
90.42
63.80
83.40
86.32
Qwen3.5-2B [ 1 ]
41.25
32.92
35.42
60.80
26.10
81.70
Granite4.2-3B [ 6 ]
85.00
78.75
82.92
65.80
76.30
92.42
Table 5: Transformer and global-attention hybrid comparisons. LCB denotes LiveCodeBench.
Benchmark
Metric
YANchor-4B
MMBench EN [ 49 ]
Answer-level accuracy
88.28
MMBench CN [ 49 ]
Answer-level accuracy
89.00
POPE [ 50 ]
Answer-level accuracy
85.43
MME [ 51 ]
Answer-level accuracy
82.98
ChartQA [ 52 ]
Answer-level accuracy
83.92
AI2D [ 53 ]
Answer-level accuracy
80.05
Table 6: YANchor-4B VL capability across six benchmarks. Scores are percentages. The mean weights the six answer-level accuracies equally.
Figure 8: Long-memory task accuracy. Left: M0, R, S, and RLVR. Right: the released model with memory disabled or enabled.
Task
All four answers correct (%)
Exact retrieval
90.63
State-update interpretation
87.50
Cross-document combination
89.06
Variable binding
53.13
All tasks
80.08
Table 7: YANchor’s complete-context success on long-memory tasks. A context is successful when all four answers are correct.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
YANchor-4B
Qwen3.5-4B [ 1 ]
GSM8K [ 25 ]
94.45
94.50
MATH-500 [ 26 ]
97.49
98.00
AIME 2024 [ 27 ]
86.72
89.58
AIME 2025 [ 27 ]
77.45
80.83
AIME 2026 [ 27 ]
84.64
87.92
HMMT Feb 2026 [ 28 ]
63.64
65.91
Appendix
Table 8: Text-capability comparison with Qwen3.5-4B. Scores are percentages.
Benchmark
Metric
Score
DROP [ 46 ]
F1
77.22
MuSR [ 47 ]
Task-macro accuracy
60.42
LogiQA2 [ 48 ]
Accuracy
72.23
Appendix
Table 9: Reading comprehension and multi-step reasoning results for YANchor-4B. All scores are percentages.
Model
GSM8K [ 25 ]
MATH [ 26 ]
AIME ’24 [ 27 ]
AIME ’25 [ 27 ]
AIME ’26 [ 27 ]
HMMT ’26 [ 28 ]
YANchor-4B
94.45
97.49
86.72
77.45
84.64
63.64
Bounded-state models
ARWKV-R1 1.5B [ 16 ]
21.00
12.60
0.00
0.00
0.00
1.89
RWKV-7 G1i 1.5B [ 14 ]
60.00
28.20
0.00
0.42
0.42
0.00
RWKV-7 G1j 1.5B [ 14 ]
57.60
31.40
1.25
2.08
0.00
0.00
RWKV-7 G1i 2.9B [ 14 ]
71.30
43.80
1.25
1.67
2.08
0.76
Appendix
Table 10: Complete mathematics results. MATH denotes MATH-500 and HMMT denotes February 2026. All scores are percentages; bold identifies YANchor-4B.
Model
MMLU [ 29 ]
Pro [ 30 ]
Redux [ 31 ]
CMMLU [ 32 ]
C-Eval [ 33 ]
GPQA [ 34 ]
Super [ 35 ]
YANchor-4B
87.87
79.60
92.00
72.40
74.47
64.61
61.69
Bounded-state models
ARWKV-R1 1.5B [ 16 ]
19.70
9.30
16.77
14.38
15.70
14.14
6.30
RWKV-7 G1i 1.5B [ 14 ]
53.10
23.10
54.82
31.56
31.80
26.77
17.20
RWKV-7 G1j 1.5B [ 14 ]
51.70
22.20
52.62
36.28
33.90
26.26
16.10
RWKV-7 G1i 2.9B [ 14 ]
65.40
40.70
62.89
33.28
34.50
26.77
22.50
Appendix
Table 11: Knowledge and science. Pro and Redux denote MMLU-Pro and MMLU-Redux; GPQA denotes GPQA-Diamond; Super denotes SuperGPQA.
Model
Wino. [ 36 ]
Hella. [ 37 ]
ARC-C [ 38 ]
ARC-E [ 38 ]
BBH [ 39 ]
BBEH [ 40 ]
YANchor-4B
77.90
64.57
95.50
98.02
85.88
29.07
Bounded-state models
ARWKV-R1 1.5B [ 16 ]
33.30
7.50
22.50
23.90
6.40
1.11
RWKV-7 G1i 1.5B [ 14 ]
46.40
38.00
68.10
83.70
23.81
4.39
RWKV-7 G1j 1.5B [ 14 ]
49.30
35.30
65.90
81.10
33.40
5.49
RWKV-7 G1i 2.9B [ 14 ]
57.10
65.00
82.00
91.90
24.68
2.82
Appendix
Table 12: Commonsense and general reasoning. Wino. denotes WinoGrande; Hella. denotes HellaSwag.
Model
HumanEval [ 41 ]
MBPP [ 42 ]
LCB v6 [ 43 ]
IFEval [ 44 ]
IFBench [ 45 ]
All 24 mean
YANchor-4B
96.72
87.81
61.63
88.03
65.29
78.64
Bounded-state models
ARWKV-R1 1.5B [ 16 ]
4.27
5.40
0.30
11.09
8.33
10.66
RWKV-7 G1i 1.5B [ 14 ]
42.07
36.60
7.40
36.97
14.67
30.40
RWKV-7 G1j 1.5B [ 14 ]
41.46
38.60
6.50
46.40
16.00
31.29
RWKV-7 G1i 2.9B [ 14 ]
58.54
49.20
11.80
43.07
17.00
37.92
Appendix
Table 13: Code and instruction following. The final column is the unweighted mean over the 24 common benchmarks.
Model
B1 decode, tokens/s
Residency
64K demand, GiB
YANchor-4B
212.98
64K
67.71
Qwen3.5-4B [ 1 ]
153.91
2K
1,061.40
Gemma 4 E4B [ 5 ]
133.75
4K
541.30
MiniCPM5-2B [ 3 ]
164.53
2K
1,351.71
Appendix
Table 14: Single-H100 endpoints. B1 decode uses 128K input and 128K output. Memory demand uses 512 independent histories at 64K context per sequence; YANchor is reconstructed from measured components, and the three baseline values are KV-growth projections. Residency gives the last fully resident measured context length.
Input
YANchor-4B
Qwen3.5-4B [ 1 ]
Gemma 4 E4B [ 5 ]
MiniCPM5-2B [ 3 ]
4K
12,799 / B512
3,103 / B64
3,940 / B128
3,152 / B64
8K
12,732 / B512
2,853 / B64
3,432 / B128
2,938 / B64
16K
12,661 / B512
2,472 / B64
2,880 / B128
2,575 / B32
32K
12,431 / B512
1,889 / B64
2,196 / B64
2,117 / B32
Appendix
Table 15: End-to-end output throughput for 32K continuations, in tokens/s. Each cell includes the batch size.
Batch
Decode, tokens/s
TTFT, s
Allocated, GiB
Reserved, GiB
32
4,025
0.35–2.89
18.84
19.02
64
6,583
0.68–5.68
20.07
21.18
128
9,119
1.36–11.38
26.67
28.92
256
11,446
2.74–22.84
39.86
43.92
320
11,170
3.46–28.34
46.45
51.36
384
12,381
4.12–34.24
53.05
58.96
Appendix
Table 16: YANchor batch scaling on one H100 80GB. Decode throughput is total decode tokens divided by total decode time over 16 length pairs. TTFT ranges describe the first batch response; memory columns are peak allocator values over the same grid.