Action-On-Item Preference Flow: A Shared Event Schema for Predictive and Generative Personalization
Organizations: KDM Lab, Dhirubhai Ambani University · Dhirubhai Ambani University · LCS2 Lab, Indian Institute of Technology Delhi
Abstract
A user's movie, news, and dialogue histories differ in their native actions and outputs, yet each interaction supplies evidence that can update user memory. We study whether these histories can train one reusable update mechanism. An action-on-item schema pairs a mapped interaction role with a content embedding, allowing shared update parameters to operate on separate user states. We establish invariance to native relabeling, bounded state changes under item-embedding perturbations, and a pooled-training bound under explicit compatibility conditions. The Multi-Timescale State Hypothesis (MTSH) specifies how this evidence enters, persists, and is consumed; PerTIDE implements it with action gating, three state-space traces, fusion, and command-conditioned readout. On PENS, the same history encoder supports both next-news prediction and personalized headline generation. In a controlled PENS-to-MovieLens experiment, a frozen source-trained core exceeds an identically structured random core by 15.23 MRR points after fitting the same target consumer. On MIND, PerTIDE retains a 4.12-point MRR advantage over a same-input three-branch state-space control. Action, readout, and trace interventions identify complementary contributions to these gains. Together, the theory and experiments support learning history updates across compatible sources and reusing them through predictive and generative consumers.
Figures & tables
| Setting | Comparator | Baseline | PerTIDE | Gain |
|---|---|---|---|---|
| MovieLens-1M movie ranking | SSD4Rec | 14.95 / 20.31 | 17.31 / 23.78 | +2.36 / +3.47 |
| MIND news ranking | MINER | 36.60 / 40.20 | 46.65 / 49.21 | +10.05 / +9.01 |
| PENS next-news ranking | LSTUR | 9.04 / 10.16 | 9.94 / 11.58 | +0.90 / +1.42 |
| PENS headline generation | DeepSeek-R1-32B (2-shot) | 0.263 / 0.174 | 0.366 / 0.355 | +0.103 / +0.181 |
| OpenAI-Reddit TL;DR generation | DeepSeek-14B (2-shot) | 0.243 / 0.109 | 0.285 / 0.294 | +0.042 / +0.185 |
| State family | Model | MRR | nDCG@5 | HR@10 |
|---|---|---|---|---|
| PAH | MeanPool-B | 29.43 | 31.22 | 33.17 |
| PAH | MaxPool-B | 28.32 | 31.16 | 32.41 |
| MDH | GRU-B | 32.64 | 35.17 | 37.92 |
| MDH | LSTM-B | 32.17 | 34.92 | 37.05 |
| LDH | D-FM Attention-B | 32.17 | 34.56 | 36.21 |
| LDH | Discretized SSM-B | 35.14 | 38.53 | 37.27 |
| 500 Target Trajectories | 3K Target Trajectories | |||||
|---|---|---|---|---|---|---|
| Method | MRR | nDCG@10 | HR@10 | MRR | nDCG@10 | HR@10 |
| UniSRec | 4.31 | 3.79 | 8.23 | 15.18 | 16.92 | 33.49 |
| RecGURU | 6.03 | 5.73 | 11.49 | 14.21 | 16.17 | 29.44 |
| PerTIDE | 6.18 | 5.94 | 12.96 | 16.61 | 19.90 | 38.75 |
| Frozen core | Head | MRR | nDCG @10 | HR@10 |
|---|---|---|---|---|
| Random | Trained | 1.115 | 0.780 | 1.44 |
| PENS-trained | Untrained | 1.371 | 0.960 | 2.18 |
| PENS-trained | Trained | 16.344 | 16.002 | 20.28 |
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Event evidence and consumer | Primary metrics |
|---|---|---|
| MovieLens-1M | Positive/negative movie events; movie ranking | MRR, nDCG@10, HR@10 |
| MIND | Click/ignore news events; news ranking | MRR, nDCG@5, HR@10 |
| PENS | Click/skip news events; next-news ranking | MRR, nDCG@10 |
| PENS | News history and headline request; personalized generation | PSE-JSD, PSE-SU4, PSE-METEOR |
| OpenAI-Reddit | Feedback-derived summary events and request; eventized TL;DR generation | PSE-JSD, PSE-SU4, PSE-METEOR |
| Symbol | Meaning | Symbol | Meaning |
| UIG and Action-On-Item Schema | |||
| User-specific interaction graph. | User-specific nodes and edges. | ||
| Initial user anchor node. | Item/content node at timestep . | ||
| Response node at timestep . | Domain-native action label. | ||
| Shared action-role vocabulary. | Schema-level action role. | ||
| Domain-specific role mapping. | Domain-specific target-locus embedding map. | ||
| Symbol | Meaning | Dimensions |
| Action-Gated Event Cells | ||
| Model / latent width (seed embeddings, memory traces, user state, task embeddings) | ||
| Action-role input width | ||
| Target-locus embedding at timestep ; induces item-side abstraction | ||
| One-hot action-role input to the gate MLP | ||
| Gate MLP layer-1 parameters ( ) | ||
| Symbol | Meaning | Dimensions |
| Command-Conditioned Readout (Realization of and ) | ||
| Task / command seed embedding | ||
| Command/task distribution ; routing uses the preceding task-ready state (Section 4 ) | ||
| Task-distribution projection | ||
| Command-to-state gating projection in | ||
| Command-gate bias | ||
| Symbol | Meaning | Dimensions |
| Decoder-UFI (User Fused Injection) | ||
| UFI MLP layer-1 parameters ( ) | ||
| UFI MLP layer-2 parameters ( ) | ||
| User-fused control, | ||
| Decoder-UGI (User Gated Injection) | ||
| FiLM-scale MLP layer-1 parameters ( ) | ||
| Symbol | Meaning | Dimensions |
| Decoder Document Grounding | ||
| Document-grounding MLP layer-1 parameters ( ) | ||
| Document-grounding MLP layer-2 parameters ( ) | ||
| Document-grounding control, | ||
| Final normalized decoder control, | ||
| Learned Prefix Injection | ||
| Metric | PerTIDE (Ours) | Single-scale encoder | Prompted LLM (ICL) |
| A. Parameter and Memory Footprint | |||
| Trainable parameters (encoder only) | 154.36M | 130–145M | 0 |
| Decoder contextualization parameters | 8.27M (UFI) / 10.63M (UGI) | task-dependent | 0 |
| Encoder + contextualization parameters | 162.63M (UFI) / 164.99M (UGI) | 130–145M + context | 0 |
| Decoder parameters fine-tuned for generation | 14.18M (last 2 blocks + final LN) | same protocol if DistilGPT2 is used | 0 (ICL) |
| Pretrained decoder backbone deployed | 82M (DistilGPT2) | 82M (if used) | 7B–32B |
| Metric | PerTIDE | Single-scale encoder | Prompted LLM |
| C. Predictive Serving Cost | |||
| Latency per sample (end-to-end) | 2–6 ms | 2–5 ms | 20–60 ms |
| Throughput (samples/sec) | 160–420 | 180–450 | 20–60 |
| FLOPs per sample | 3.8 | 2.5–3.2 | 1 |
| D. Generative Serving Cost | |||
| Contextualization | UFI / UGI learned prefix | concat / none | ICL prompt |
| Regime | Model | MRR | nDCG@10 | HR@10 |
|---|---|---|---|---|
| PAH | BPR-MF | 6.50 | 6.12 | 12.81 |
| MDH | GRU4Rec | 12.76 | 16.42 | 29.01 |
| Caser | 13.54 | 19.21 | 28.92 | |
| Diff4Rec | 12.02 | 15.72 | 22.38 | |
| LDH | S 3 Rec | 10.62 | 13.67 | 19.21 |
| SASRec | 13.20 | 19.96 | 30.23 |
| Regime | Model | MRR | nDCG@5 | HR@10 |
|---|---|---|---|---|
| PAH | NAML | 32.75 | 35.66 | 41.40 |
| PLM-NR | 35.39 | 38.71 | 44.38 | |
| Mean Pooling-B (ours) | 29.43 | 31.22 | 33.17 | |
| Max Pooling-B (ours) | 28.32 | 31.16 | 32.41 | |
| MDH | EBNR | 31.26 | 32.18 | 39.04 |
| GRU-B (ours) | 32.64 | 35.17 | 37.92 |
| Regime | Model | MRR | nDCG@5 | nDCG@10 | HR@10 |
|---|---|---|---|---|---|
| MDH | EBNR | 2.65 | 1.82 | 2.45 | 4.87 |
| PAH | NAML | 1.29 | 0.43 | 0.81 | 2.01 |
| LDH | NRMS | 1.18 | 0.39 | 0.74 | 1.92 |
| LDH | TrRMIo | 8.02 | 7.85 | 8.53 | 12.77 |
| LDH | LSTUR | 9.04 | 8.21 | 10.16 | 12.93 |
| LDH | MINER | 5.17 | 4.71 | 5.61 | 9.77 |
| Dataset | Family | Model / Variant | L-Tr | S-Tr | E-Tr |
|---|---|---|---|---|---|
| MovieLens | Baseline | Mamba4Rec | 26.43/29.38 | 26.12/28.87 | 25.72/28.31 |
| GRU4Rec | 24.17/27.32 | 23.65/26.58 | 22.13/24.71 | ||
| SASRec | 25.82/28.16 | 25.85/28.45 | 25.17/27.91 | ||
| PerTIDE | PerTIDE-L | 27.31/29.96 | 27.18/29.35 | 25.83/28.77 | |
| PerTIDE-S | 25.21/28.64 | 27.13/29.22 | 22.16/24.84 | ||
| PerTIDE-E | 22.19/24.08 | 22.75/25.04 | 27.14/30.36 |
| Variant | MIND Recommendation | ML-1M Recommendation | PENS Summarization | ||||||
|---|---|---|---|---|---|---|---|---|---|
| MRR | nDCG@5 | HR@10 | MRR | nDCG@10 | HR@10 | PSE-JSD | PSE-SU4 | PSE-METEOR | |
| L only | 40.41 | 43.75 | 46.13 | 12.31 | 16.15 | 20.34 | 0.231 | 0.095 | 0.107 |
| S only | 37.65 | 41.32 | 44.24 | 12.47 | 16.62 | 20.81 | 0.238 | 0.106 | 0.111 |
| E only | 24.18 | 27.73 | 31.23 | 10.18 | 13.17 | 18.54 | 0.076 | 0.042 | 0.045 |
| L + S | 41.26 | 43.95 | 46.72 | 14.17 | 18.44 | 22.11 | 0.264 | 0.143 | 0.182 |
| L + E | 40.45 | 43.82 | 46.44 | 14.03 | 18.07 | 22.74 | 0.246 | 0.103 | 0.108 |
| Category | Model | PSE-JSD | PSE-SU4 | PSE-METEOR |
|---|---|---|---|---|
| LLMs (2-shot) | LLaMA-13B | 0.227 | 0.078 | 0.081 |
| DeepSeek-14B | 0.248 | 0.094 | 0.097 | |
| Gemini-2.5-Flash | 0.222 | 0.104 | 0.124 | |
| DeepSeek-R1-32B | 0.263 | 0.125 | 0.174 | |
| Qwen-2.5-32B | 0.162 | 0.103 | 0.115 | |
| Prompt-Chaining | Mistral-7B | 0.072 | 0.026 | 0.023 |
| Category | Model | RG-SU4 | BLEU-4 | BScore | HJ |
|---|---|---|---|---|---|
| Specialized (Personalized) | PENS-NRMS-T2 | 13.64 | 4.48 | 86.13 | 2.95 |
| GTP-TrRMIo | 21.91 | 10.31 | 88.53 | 2.44 | |
| SP-Individual | 19.54 | 8.90 | 86.61 | 2.86 | |
| LLMs (2-shot history) | LLaMA-13B | 18.31 | 11.85 | 88.76 | 3.05 |
| DeepSeek-14B | 19.57 | 12.68 | 89.43 | 3.03 | |
| DeepSeek-32B | 29.42 | 19.31 | 90.87 | 3.12 |
| Relation | Spearman | User-clustered 95% CI |
|---|---|---|
| vs. frozen risk | +0.109 | [+0.002,+0.207] |
| vs. adapted risk | +0.160 | [+0.060,+0.258] |
| vs. frozen RR gain | -0.089 | [-0.169,-0.004] |
| vs. adapted RR gain | -0.114 | [-0.195,-0.029] |
| Role flips | MRR | nDCG@10 | HR@10 |
|---|---|---|---|
| 0% | 12.14 | 16.02 | 32.0 |
| 10% | 10.13 | 13.02 | 25.0 |
| 25% | 9.17 | 11.10 | 22.0 |
| 50% | 7.71 | 8.93 | 15.0 |
| 100% | 2.81 | 4.90 | 10.0 |
| Events | MRR | nDCG@10 | HR@10 |
|---|---|---|---|
| 10 | 11.874 | 14.904 | 27.8 |
| 20 | 11.881 | 15.011 | 28.2 |
| 30 | 11.909 | 15.091 | 28.4 |
| 40 | 11.926 | 15.157 | 28.6 |
| Category | Model | PSE-JSD | PSE-SU4 | PSE-METEOR |
|---|---|---|---|---|
| LLMs (2-shot) | LLaMA-13B | 0.232 | 0.093 | 0.107 |
| Zephyr-7B | 0.214 | 0.087 | 0.104 | |
| Mistral-7B | 0.226 | 0.088 | 0.103 | |
| DeepSeek-14B | 0.243 | 0.095 | 0.109 | |
| MTSH | PerTIDE -Full | 0.285 | 0.261 | 0.294 |
| Model | PSE-JSD | PSE-SU4 | PSE-METEOR |
|---|---|---|---|
| DeepSeek-R1-32B (2-shot) | (−42.21%) | (−30.4%) | (−47.2%) |
| Gemini-2.5-Flash (2-shot) | (−45.0%) | (−41.3%) | (−43.5%) |
| Best Baseline (GTP) | (−33.3%) | (−47.1%) | (−42.1%) |
| PerTIDE | (−42.35%) | (−28.45%) | (−26.2%) |
| Category | Model | RG-2 | RG-L | BLEU-4 | BScore |
|---|---|---|---|---|---|
| LLMs (2-shot) | DeepSeek-32B | 9.71 | 8.46 | 5.61 | 71.42 |
| Qwen-2.5-32B | 9.43 | 8.21 | 4.75 | 68.83 | |
| MS/Phi-Instruct | 9.49 | 8.05 | 3.21 | 69.52 | |
| PAH | Meanpool-B | ||||
| Maxpool-B | |||||
| MDH | GRU-B |
| Model | R@1 | MRR |
|---|---|---|
| SBERT (zero-shot) | 35.67 | 45.75 |
| SBERT | ||
| SBERT+ViT ( ) | ||
| SBERT+ViT ( ) | ||
| SBERT+ViT ( ) | ||
| SBERT+ViT ( , Full) |
| Category | Model | R@1 | MRR |
|---|---|---|---|
| Text Only | SBERT ( ) | ||
| SBERT+ViT | |||
| (Full) | |||
| SBERT+CLIP | |||