Personalization encoders compress evolving interaction histories into preference states used to rank items or condition text generation. A task head operating only on this state can miss useful evidence that remains in the frozen encoder's cached representations for individual timesteps. We study this recoverability gap and propose REPAIR, which compares cached representations with the current preference state in a compact learned coordinate space. It resolves corrective evidence over extended history, recent interactions, and localized bursts. It then selects which patterns at which timesteps contribute and adds their aggregate correction to the state before the task head. Encoder-host repair reuses representations from the existing forward computation without re-encoding the history. Across MovieLens, PENS, MIND, and Amazon Reviews 2023, training only REPAIR improves MRR and nDCG@10 for all twelve representative recommendation hosts while both encoder and task head remain frozen. Head-only finetuning of the same hosts yields smaller gains. For example, Mamba4Rec on MovieLens gains 3.96 MRR points, compared with 0.19 from head-only finetuning. Rank and temporal diagnostics support a compact, host-dependent corrective structure. In personalized generation, IMPerSumm improves the two reported weighted PerSEval variants, which assess responsiveness to user preference, by up to 25.23%. These results support post-compression state correction and distinguish the availability of preference evidence from its downstream use.
Figures & tables
Figure 1: Overview of REPAIR . Given a frozen preference state and cached per-timestep representations from host computation, REPAIR constructs state-relative corrective evidence, resolves it through long-, short-, and episodic-support branches, selectively retains timestep-coordinate corrective components, and fuses the aggregated correction with the frozen state before the downstream interface consumes the repaired state.
Task
Dataset
Host Coverage
Test Instances
Movie recommendation
MovieLens
9 sequential hosts
15K
News recommendation
MIND
7 history encoders
20K
News recommendation
PENS
History encoders and LLMs
20K
Product recommendation
Amazon Reviews 2023
3 multimodal hosts
2K
Headline generation
PENS
11 encoder interfaces and 2 LLMs
20K
Table 1: Evaluation design. Recommendation ranks one target against 499 negatives. Host architectures and citations are in Appendix E , with protocols in Appendix C .
Dataset
Host
MRR / nDCG@10
Original Host
Head-only FT
REPAIR (Frozen Head)
REPAIR + Head Alignment
MovieLens
S: TiSASRec
35.51 / 41.14
35.57 / 41.19
36.38 / 42.07
36.78 / 42.91
M: Mamba4Rec
29.29 / 38.50
29.48 / 38.89
33.25 / 39.20
34.26 / 40.17
W: SASRec
26.43 / 33.82
27.05 / 34.37
29.57 / 40.08
28.49 / 35.18
PENS
S: LSTUR
10.44 / 12.15
10.98 / 12.32
11.98 / 12.71
12.85 / 15.88
M: EBNR
2.65 / 2.45
2.93 / 2.86
4.07 / 3.97
11.97 / 14.28
Table 2: RQ1: Post-compression recoverability under frozen-host controls. Columns compare the original host, head-only continuation on the unchanged state, REPAIR with encoder and head frozen, and two-epoch joint REPAIR -head alignment with the host encoder frozen. Values are MRR / nDCG@10. Three-seed means, variability, and bootstrap intervals are available for OpenCLIP and SigLIP2 on Amazon and EBNR on PENS (Appendix Table 10 ). The remaining rows are reported as point estimates. S/M/W denote strongest, median, and weakest original hosts.
Host / Dataset
Original Host
REPAIR Rank K
32
64
128
192
OpenCLIP / Amazon
1.59 / 1.05
9.31 / 9.78
11.17 / 12.22
19.49 / 22.08
11.63 / 12.16
SigLIP2 / Amazon
2.40 / 2.18
10.44 / 11.36
8.42 / 8.77
17.72 / 20.35
13.46 / 13.15
EBNR / PENS
2.65 / 2.45
6.13 / 6.19
8.65 / 8.58
11.97 / 14.28
11.23 / 13.17
Mamba4Rec / MovieLens
29.29 / 38.50
31.47 / 38.85
32.38 / 39.26
34.26 / 40.17
34.02 / 39.96
LSTUR / MIND
51.43 / 59.63
52.67 / 62.14
53.49 / 64.13
55.84 / 66.92
54.73 / 66.12
Table 3: RQ2: Repair-space rank sensitivity after compression. MRR / nDCG@10 for two-epoch downstream alignment across K . Bold marks the strongest repair rank for each host.
Controlled one-view diagnostic
Full REPAIR ( L+S+E jointly)
Dataset
Host
L-Tr
S-Tr
E-Tr
L-Tr
S-Tr
E-Tr
Copy-through prediction
L
S
E
all three views used jointly
MovieLens
Mamba4Rec
S [33.95]
S [33.92]
S [33.88]
[34.87]
[35.52]
[34.39]
TiSASRec
L [36.19]
S [36.20]
S [36.04]
[37.38]
[36.93]
[36.63]
MIND
NAML
L [52.12]
S [52.03]
L [53.70]
[53.31]
[54.55]
[53.80]
EBNR
S [55.08]
L [53.64]
E [53.70]
[56.79]
[55.69]
[53.19]
Table 4: RQ2 temporal diagnostic. If the temporal module merely copied the controlled input timing, the copy-through prediction would be L , S , and E for L-Tr, S-Tr, and E-Tr, respectively. This pattern holds in only 6/12 settings. Full REPAIR , which jointly uses L+S+E , outperforms the best isolated view in 11/12 settings. Brackets report MRR.
Tasks (Datasets)
Host Model
Base
+ SERAC-Corr
+ Rec-Denoiser-Corr
+ REPAIR
MRR
nDCG@10
MRR
nDCG@10
MRR
nDCG@10
MRR
nDCG@10
Movie Reco. (MovieLens)
TiSASRec
35.51
41.14
33.17
41.32
34.21
42.08
36.78
42.91
Mamba4Rec
29.29
38.50
29.81
38.63
31.32
39.22
34.26
40.17
SIGMA
33.27
38.96
31.95
38.98
29.43
39.61
34.81
40.72
News Reco. (MIND)
LSTUR
51.43
59.63
53.66
62.17
52.45
61.35
55.84
66.92
NAML
45.65
46.56
46.47
51.33
48.12
48.94
53.22
63.18
Table 5: RQ3: Intervention-locus comparison under the same frozen-host protocol. Pre-head REPAIR is compared with two adapted post-head correctors using the same task-trained host and two-epoch downstream alignment protocol.
Host
RG-L
METEOR
PSE-W-JSD
PSE-W-RG-L
Walk2Pers
+6.58
+13.38
+13.25
+15.32
IMPerSumm
-2.69
+4.76
+20.76
+25.23
DeepSeek-32B
+53.21
+45.49
+9.32
+5.05
Qwen2.5-32B
+41.58
+47.64
+31.52
+12.61
Table 6: RQ3 generation and personalization on PENS. Values are relative gains (%). Full ranking and generation comparisons are in Appendix Table 17 . Absolute scores are in Appendix Tables 15 and 16 .
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2: Encoder-host repair from cached evidence to the task head. Read from left to right. The state and cached sequence enter a shared space, the low-rank reader produces corrective coordinates, and centering precedes temporal resolution and coordinate-wise selection. State gating and weighted aggregation form a correction, which is fused with the original state. The diagram’s raw-residual label denotes r(tj;ti) in Appendix A.3 . Dashed arrows indicate conditioning, and the original state remains the residual base.
Figure 3: Prompt-conditioned repair for a frozen LLM. The prompt computation supplies the terminal state, while a separate interaction construction supplies corrective evidence. Repair maps the fused state into continuous prefix vectors. The frozen decoder then evaluates the prefix-conditioned input to produce the task output. Decoder weights remain fixed, and the attention cache must correspond to this modified input.
Figure 4: Task prompts used for the LLM headline-generation and recommendation settings. The history supplies preference evidence, while the task block specifies the requested output and, for generation, the current source document. The examples illustrate prompt structure rather than additional evaluation instances. The repair-specific continuous prefix is constructed separately as described in Appendix A.7 .
Symbol
Meaning
Symbol
Meaning
Interaction and host interface
Gu=⟨Nu,Eu⟩
User interaction graph
Nu
User-specific typed node set
Eu
Action-labeled edge set
u(t0)
Initial user node
v(ti)
Content or response node
a(ti)
Observed action label
{pos,neg,cmd}
Normalized action types
b(ti)=⟨a(ti),v(ti)⟩
Interaction unit at timestep
τbt1:i
Model-side prefix trajectory
ti
Current prefix endpoint
Appendix
Table 7: Notation for the interaction sequence, frozen host interface, repair operators, and supervision. Hats on Urep and Vrep denote column normalization. The family weights denoted by ω in the main text are denoted by λ here, with ω reserved for within-family coefficients.
Symbol
Meaning
Symbol
Meaning
Temporal resolution
ρ∈{L,S,E}
Temporal basis index
κL(ℓj)
Long-support temporal basis
κS(ℓj)
Short-support temporal basis
κE(ℓj)
Episodic-support temporal basis
β
Learned long-support decay vector
ξ
Learned episodic-support rate vector
WL,WS,WE
Temporal branch-mixing matrices
Grep
Multi-basis resolution operator
η(tj)
Temporally resolved corrective vector
Appendix
Table 8: Notation continued from Table 7 . Temporal resolution, selective retention, aggregation, and training objective.
Component
Symbol / Configuration
Dimension / Value
Core repair-interface dimensions
Repair-interface state dimension
dz
768
Repair-interface cached-representation dimension
de
768
Shared repair space
m
192
Repair rank
K
128
Maximum history length
L
50
Appendix
Table 9: Architecture and fixed hyperparameters for the encoder-host repair interface. The dimensions describe the implemented common interface, not every host’s native hidden width. The default K=128 is varied only in the rank sweep. Appendix C.4.4 defines the loss families and within-family coefficients.
Setting
Hbest
R
RH2
OpenCLIP / Amazon
2.21±0.18 [2.04, 2.36]
4.43±0.25 [4.17, 4.64]
19.49±1.43 [18.97, 20.06]
SigLIP2 / Amazon
2.67±0.14 [2.33, 2.91]
4.16±0.17 [3.97, 4.33]
17.72±2.34 [17.37, 18.04]
EBNR / PENS
2.93±0.31 [2.73, 3.05]
4.07±0.19 [3.98, 4.24]
11.97±0.62 [11.76, 12.25]
Appendix
Table 10: Repeated-run reliability for three RQ1 settings. Entries are MRR mean ± standard deviation across three training seeds, followed by 95% user-clustered bootstrap intervals from 10,000 resamples. MRR uses the 0 to 100 scale.
Setting
Host
MRR
nDCG@5
Base
+ REPAIR
Δ
Base
+ REPAIR
Δ
Direct
Mamba4Rec
29.29
34.26
+4.97
30.07
35.56
+5.49
TiSASRec
35.51
36.78
+1.27
36.87
38.41
+1.54
HSTU
29.00
30.62
+1.62
29.69
31.55
+1.86
FEARec
28.03
30.62
+2.59
28.20
31.55
+3.35
SIGMA
33.27
34.81
+1.54
34.43
36.26
+1.83
Appendix
Table 11: MovieLens recommendation under Stage-3 repair and head alignment, with the host encoder frozen. Base is the original task-trained host. Direct rows train repair for that host. Transfer rows use the module trained on Mamba4Rec. Scores are on a 0 to 100 scale and higher is better. Each Δ is an absolute score-point change. All paired conditions rank the same target against 499 negatives. Appendix G.1 distinguishes these aligned results from frozen-head attribution.
Host
MRR
nDCG@5
Base
+ REPAIR
Δ
Base
+ REPAIR
Δ
EBNR
2.65
11.97
+9.32
1.82
12.68
+10.86
NAML
1.29
12.38
+11.09
0.43
13.16
+12.73
NRMS
1.18
12.11
+10.93
0.39
12.89
+12.50
TrRMIo
8.02
12.86
+4.84
7.85
13.97
+6.12
SCAPE
1.57
8.34
+6.77
0.76
7.91
+7.15
Appendix
Table 12: PENS recommendation with paired target and 499 -negative candidate sets. Base denotes the task-trained host. Repaired encoder rows use Stage-3 repair and head alignment with frozen encoders. LLM rows use the frozen-decoder soft-prefix interface in Appendix A.7 . Scores are on a 0 to 100 scale, higher is better, and Δ denotes absolute score-point change.
Host
MRR
nDCG@5
Base
+ REPAIR
Δ
Base
+ REPAIR
Δ
EBNR
47.08
53.22
+6.14
48.65
57.21
+8.56
NAML
45.65
53.22
+7.57
43.13
57.21
+14.08
NRMS
44.91
52.74
+7.83
42.38
56.48
+14.10
TrRMIo
47.43
55.21
+7.78
46.27
52.83
+6.56
SCAPE-GRU
42.53
45.32
+2.79
37.53
44.15
+6.62
Appendix
Table 13: MIND recommendation under Stage-3 repair and head alignment with frozen encoders. Base denotes the original task-trained host. Paired conditions rank the same target against 499 negatives. Scores are on a 0 to 100 scale, higher is better, and Δ denotes absolute score-point change. Appendix G.1 discusses the seven hosts.
Host
MRR
nDCG@5
Base
+ REPAIR
Δ
Base
+ REPAIR
Δ
SigLIP2
2.40
17.72
+15.33
1.53
17.71
+16.18
OpenCLIP
1.59
19.49
+17.91
0.64
19.46
+18.83
DINOv2 E5
1.19
16.15
+14.96
0.43
16.30
+15.87
Appendix
Table 14: Amazon Reviews 2023 All_Beauty recommendation with frozen multimodal encoders and Stage-3 repair and head alignment. Each instance ranks one target against the same 499 negatives in paired conditions. Scores are on a 0 to 100 scale, higher is better, and Δ is an absolute score-point change.
Host
RG-1
RG-L
RG-SU4
Base
+ REPAIR
Gain (%)
Base
+ REPAIR
Gain (%)
Base
+ REPAIR
Gain (%)
DeepSeek-32B
20.62
26.84
+30.16%
13.85
21.22
+53.21%
4.42
11.99
+171.27%
Qwen2.5-32B
18.51
23.47
+26.80%
12.53
17.74
+41.58%
3.86
8.46
+119.17%
FPG
16.48
19.51
+18.39%
13.08
14.12
+7.95%
2.51
3.11
+23.90%
SCAPE-GRU
19.02
22.47
+18.14%
15.26
16.41
+7.54%
2.88
3.75
+30.21%
GTP
15.97
18.28
+14.46%
12.75
13.67
+7.22%
2.45
2.92
+19.18%
Appendix
Table 15: Full PENS headline-generation quality results. Base and repaired values equal 100 times the native metric score. Higher is better. Gain is the reported relative percentage change from Base. RG abbreviates ROUGE, W denotes weighted overlap, and P denotes paraphrase-aware overlap. Both panels use the same generated outputs. Encoder hosts use the aligned repair condition, while LLMs use the frozen soft-prefix interface. Definitions are in Appendix D.3 .
Host
PSE-W-JSD
PSE-W-RG-L
Base
+ REPAIR
Gain (%)
Base
+ REPAIR
Gain (%)
DeepSeek-32B
118
129
+9.32%
148.5
156
+5.05%
Qwen2.5-32B
92
121
+31.52%
125.3
141.1
+12.61%
SCAPE
103
112.5
+9.22%
128
135.2
+5.62%
FPG
74.8
83.9
+12.17%
104.2
110.4
+5.95%
GTP
89.5
98.7
+10.28%
113.8
120.8
+6.15%
Appendix
Table 16: Full PENS personalization-sensitive results on the same outputs as Table 15 . Base and repaired scores are displayed in units of 10−3 , so a displayed value of 118 denotes a native score of 0.118 . Higher is better, and Gain is the reported relative percentage change. PSE denotes PerSEval. JSD, W, P, and RG are defined in Appendix D.4 . Encoder and LLM repair interfaces follow the same distinction as in the generation-quality table.
Host
Interface
Ranking
Generation Quality
Personalization
MRR
nDCG@10
RG-L
METEOR
PSE-W-JSD
PSE-W-RG-L
NAML
T1
+859.69
+1767.90
+8.20
+25.24
+32.66
+9.94
T2
+859.69
+1767.90
+7.96
+23.28
+32.30
+9.49
EBNR
T1
+351.70
+482.86
+7.76
+27.51
+31.09
+10.33
T2
+351.70
+482.86
+7.48
+26.87
+30.25
+9.53
NRMS
T1
+926.27
+1916.22
+8.21
+24.14
+35.15
+10.20
Appendix
Table 17: Full RQ3 comparison of task-specific downstream use on PENS. Normalized relative gains (%) after preference-state repair across ranking, generation-quality, and personalization-sensitive metrics. The original results in absolute values are in Appendix Tables 12 , 15 , and 16 .
Dataset
Host
Base
NoTS
L
S
E
L+S
S+E
L+E
L+S+E
MovieLens
Mamba4Rec
29.29
28.43
29.92
29.63
29.45
32.17
31.35
31.04
34.26
TiSASRec
35.51
35.52
35.54
35.78
35.83
35.91
35.66
35.41
36.78
SIGMA
33.27
33.29
33.66
33.32
33.29
33.88
33.73
33.43
34.81
MIND
NAML
45.65
45.95
47.65
48.32
47.13
49.31
49.18
48.79
53.22
EBNR
47.08
47.74
48.88
48.03
47.62
50.13
49.38
48.76
53.22
NRMS
44.91
45.06
46.43
45.13
47.73
47.58
46.44
46.15
52.74
Appendix
Table 18: Temporal-pathway ablation on MovieLens and MIND. All values are MRR on a 0 to 100 scale, with higher values better. Base is the original host, NoTS removes temporal resolution, and L , S , and E retain the indicated long-support, short-support, and episodic-support branches. The full configuration is L+S+E . Mamba4Rec and SIGMA are state-space hosts, TiSASRec and NRMS use self-attention, NAML uses attentive aggregation, and EBNR is recurrent. These aligned configurations are distinct from the controlled input-timing diagnostic in Table 4 .
Dataset
Host
Base
Ltask
Ltask+Lpos
Ltask+Lpos+Lcoh
MovieLens
TiSASRec
35.51
35.71
36.14
36.78
Mamba4Rec
29.29
31.43
32.16
34.26
MIND
LSTUR
51.43
52.17
54.65
55.84
NRMS
44.91
46.39
51.29
52.74
Appendix
Table 19: Progressive supervision ablation on MovieLens and MIND. Values are MRR on a 0 to 100 scale, with higher values better. Successive columns add task utility, timestep provenance, and source coherence. This fixed-order comparison measures the evaluated combinations and does not isolate each family under every possible combination. Appendix G.4 connects it to the conditional local theorem.
Host
i
MRR
nDCG@10
SelSteps
Base
NS
SEL
Base
NS
SEL
Mamba4Rec
10
17.43
19.32
19.45
22.62
27.15
27.94
7
20
21.65
23.22
23.43
27.85
30.61
31.24
16
30
26.11
28.45
28.29
33.32
35.40
35.36
22
40
27.32
29.89
29.45
35.53
36.41
36.21
28
50
29.29
33.95
34.26
38.50
40.03
40.17
34
Appendix
Table 20: Length-controlled selective-retention ablation. At each history length i , Base is the original host, NS uses non-selective repair, and SEL uses adaptive timestep-coordinate retention. MRR and nDCG@10 use a 0 to 100 scale, with higher values better. SelSteps counts positions retained by at least one corrective coordinate. Mamba4Rec and TiSASRec use MovieLens. NAML and EBNR use MIND. Comparisons between NS and SEL hold the prefix length fixed. Appendix G.5 discusses both gains and reversals.
Dataset
Host Encoder
Noisy Base
+ REPAIR
MRR
nDCG@10
MRR
nDCG@10
MovieLens
Mamba4Rec
18.21
23.94
21.43
25.74
MovieLens
TiSASRec
22.34
25.92
24.65
28.17
MIND
NAML
28.61
29.23
32.85
35.81
MIND
EBNR
29.63
32.44
31.65
35.17
Appendix
Table 21: Noisy-history stress test with approximately 35 to 40% of valid positions replaced with random training-pool content. The supervised target and the relative order of non-corrupted interactions are preserved. Noisy Base and repaired scores use the same corrupted input without retraining. MRR and nDCG@10 are on a 0 to 100 scale, with higher values better. Appendix G.6 states the scope of the test.
Host
Base MRR
+ REPAIR
Δ
Gain (%)
Mamba4Rec
29.29
34.26
+4.97
16.96
TiSASRec
35.51
36.78
+1.27
3.57
EBNR
47.08
53.22
+6.14
13.04
NAML
45.65
53.22
+7.57
16.58
DeepSeek-32B
1.31
1.83
+0.52
39.69
Appendix
Table 22: Serving overhead relative to the same host in the same execution environment. Base, repaired MRR, and Δ use a 0 to 100 scale. Gain is a relative percentage. Memory overhead is derived from parameter counts and excludes activation and cache residency. Throughput change is derived from reciprocal latency. The lower panel divides multiplicative score gain by multiplicative resource cost. Ratios above one favor score gain relative to cost. Ratios below one favor cost growth. The LLM row uses its separate soft-prefix serving protocol.
Model
Work Unit
Base
Added Repair
Total
Mamba4Rec
GFLOPs/query
0.75
0.18
0.93
TiSASRec
GFLOPs/query
1.45
0.18
1.63
NAML
GFLOPs/query
0.03
0.02
0.05
DeepSeek-32B
GFLOPs/token
64.0
unreported
64.0
Appendix
Table 23: Arithmetic work for representative serving paths at history length T=50 . GFLOPs denotes billions of floating-point operations. Budget columns are required arithmetic throughput, computed as work divided by 50 or 25 milliseconds. They are not measured latencies. Scaling describes the host path. Encoder work is per query, whereas DeepSeek work is per generated token. For DeepSeek, added repair work is unreported, and the total is therefore a decoder reference rather than a complete repaired-system estimate.
attention score/value passes + repair fusion. Latency is more sensitive to sequence interactions because the host is attention-based rather than SSM-based.
NAML . cached news-table serving
Frozen : cached NEWS_TABLE + frozen user encoder. Added : repair module attached after cached-history composition.
table lookup / memory bandwidth + attentive user aggregation. Serving shifts a large fraction of work out of online text encoding and into cached vector retrieval.
token-by-token decoding. Prefix-conditioned decoding builds the appropriate attention cache for the modified input.
Appendix
Table 24: Resident artifacts and online operations for the evaluated serving paths. A within-forward sequence cache, a persistent news-vector table, and an attention key-value (KV) cache serve different roles. The LLM decoder is evaluated with the inserted soft prefix, so a KV cache for an unmodified prompt cannot be reused unchanged. Appendix G.7 relates these paths to the measured and estimated costs.
Host
MRR
Hit@10
Base
+ REPAIR
Δ
Base
+ REPAIR
Δ
SASRec
4.11
4.56
+0.45
7.82
8.76
+0.94
TextCNN
2.57
2.78
+0.21
5.08
5.04
-0.04
GRU4Rec
4.14
4.56
+0.42
8.80
9.13
+0.33
Appendix
Table 25: TG-ReDial recommendation and response generation from dialogue-derived interactions. Base and repaired conditions use the same held-out task instances. Recommendation ranks one target against 499 negatives. Reported scores are on a 0 to 100 scale, with higher values better. Δ is an absolute score-point change. The first two panels evaluate recommendation, and the third evaluates generation. RG denotes ROUGE. The panels use distinct task interfaces. We discuss the mixed TextCNN response and the difference between the two generators in Appendix G.8 .
A user's movie, news, and dialogue histories differ in their native actions and outputs, yet each interaction supplies evidence that can update user memory. We study whether these histories can train one reusable update mechanism. An action-on-item schema pairs a mapped interaction role with a content embedding, allowing shared update parameters to operate on separate user states. We establish invariance to native relabeling, bounded state changes under item-embedding perturbations, and a pooled-training bound under explicit compatibility conditions. The Multi-Timescale State Hypothesis (MTSH) specifies how this evidence enters, persists, and is consumed; PerTIDE implements it with action gating, three state-space traces, fusion, and command-conditioned readout. On PENS, the same history encoder supports both next-news prediction and personalized headline generation. In a controlled PENS-to-MovieLens experiment, a frozen source-trained core exceeds an identically structured random core by 15.23 MRR points after fitting the same target consumer. On MIND, PerTIDE retains a 4.12-point MRR advantage over a same-input three-branch state-space control. Action, readout, and trace interventions identify complementary contributions to these gains. Together, the theory and experiments support learning history updates across compatible sources and reusing them through predictive and generative consumers.
Sequential recommender systems typically infer user preferences through single-pass encoding of interaction histories without iterative refinement, relying on increasingly deep architectures to capture complex patterns. In this work, we revisit sequential recommendation from a recursive inference perspective: can user preferences be modeled as a persistent latent state that is recursively refined? We propose RecRec (Recursive Recommendation), a lightweight model that maintains a compact latent state and updates it through a shared recursive module conditioned on interaction evidence. Unlike prior recursive models, RecRec introduces an evidence-anchored correction mechanism that stabilizes refinement by grounding each update in the original interaction context, preventing semantic drift during deep recursive reasoning. Experiments on three benchmark datasets under standard evaluation protocols show that RecRec matches or outperforms state-of-the-art sequential, graph-based, and reasoning-enhanced recommenders while using only 3.9M to 14M parameters. Ablation studies demonstrate that both recursive refinement and the evidence-anchored correction gate contribute significantly to performance, highlighting the effectiveness of recursive latent inference as a scalable alternative to deeper or language-based architectures. Code is available at https://anonymous.4open.science/r/RecRec-6B67/README.md.
Personalized generation with frozen large language models requires a conditioning signal that is both compact and current. Existing personalization methods typically retrieve or summarize user histories in text, or compress them into static latent profiles and soft prompts. These approaches are efficient, but they treat a user's past behavior as an aggregate profile and therefore mix stable identity, recent drift, and item content in the same representation. We propose LAtent Trajectory Tracking and Extrapolation (LATTE), a framework that represents personalization as forecasting a peer anchored relative preference state. For each historical session, LATTE subtracts a time masked baseline formed from comparable users who responded to the same item, producing a state that measures how the target user differs from peers under a shared item context. A lightweight sequence predictor then forecasts the next state in this trajectory, and a State to Token Bridge injects the forecast into a frozen instruction tuned LLM through a single anchored soft token. We provide a latent factor analysis showing when peer anchoring cancels shared item variation and why temporal forecasting trades off stale averages against noisy recent states. Experiments on Amazon Reviews 2023 and MemoryCD show that LATTE consistently outperforms retrieval, summary memory, static latent profiles, difference aware latent profiles, and soft prompt compression baselines. On Amazon Reviews 2023, LATTE improves average ROUGE-L from 0.219 for a static latent profile and 0.245 for the strongest added latent compression baseline to 0.259. Additional pairwise comparisons and diagnostic analyses suggest that the improvement is mainly due to forecasting user-specific trajectory information, rather than merely adding a soft prompt interface.
Jinze Li, Xiaoyan Yang, Shuo Yang +5
The University of Hong Kong · Ant Healthcare, Ant Group