RouteRec: Behavior-Guided Sparse Routing for Sequential Recommendation
Organizations: Korea Advanced Institute of Science and Technology Daejeon, Republic of Korea · Seoul National University Seoul, Republic of Korea
Abstract
Sessionized interaction histories contain behavioral patterns that can improve sequential recommendation. However, existing models process all sessions through the same parameterized blocks, regardless of their behavioral differences. Mixture of Experts (MoE) enables conditional computation, but it leaves open what should guide expert allocation. We propose RouteRec, a sequential recommender that uses observed session behavior as the routing criterion. RouteRec summarizes four types of behavioral evidence from sessionized histories: interaction tempo, item-group focus, repetition and carryover, and popularity tendency. It uses these cues to route computation at macro, mid, and micro scopes. Cue-derived scores first select expert groups; within each selected group, the current backbone state then refines expert selection. Across six public datasets and 18 dataset-metric combinations, RouteRec ranks first in 12 and second in three, yielding the best overall average rank of 1.61 compared with 4.11 for the next-best baseline. Additional analyses suggest that the behavioral cues guide expert allocation beyond added capacity and produce routing patterns aligned with observed behavior. Our code is available at https://github.com/jy1559/RouteRec
Figures & tables
| Family | Macro examples | Mid examples | Micro examples |
|---|---|---|---|
| Tempo | session gaps, pace trend | interval mean/std | last gap, local pace shift |
| Focus | theme consistency | category concentration, switching | local switch, suffix focus |
| Memory | repetition, carryover | repeat rate, novelty | recent reuse, longest run |
| Popularity | popularity level, drift | mean popularity, spread, trend | last-item pop., suffix spread, pop. delta |
| Dataset | Metric | SASRec | GRU4Rec | TiSASRec | FEARec | DuoRec | BSARec | FAME | DIF-SR | FDSA | RouteRec |
| KuaiRec | HR@10 | .4612 .0028 | .4352 .0108 | .3299 .0201 | .3653 .0309 | .4779 .0107 | .4174 .0098 | .4809 .0017 | .3201 .0655 | .4671 .0092 | .4836 .0035 |
| NDCG@10 | .4319 .0028 | .4020 .0115 | .2682 .0200 | .3101 .0419 | .4555 .0168 | .3739 .0155 | .4617 .0015 | .2802 .0887 | .4377 .0137 | .4656 .0041 | |
| MRR@10 | .4229 .0029 | .3918 .0117 | .2491 .0199 | .2932 .0452 | .4486 .0187 | .3604 .0171 | .4558 .0016 | .2679 .0959 | .4286 .0151 | .4602 .0043 | |
| LastFM | HR@10 | .3439 .0018 | .3281 .0017 | .3353 .0011 | .2978 .0022 | .3447 .0010 | .2278 .0013 | .2465 .0011 | .2884 .0014 | .3321 .0014 | .3569 .0006 |
| NDCG@10 | .2703 .0018 | .2645 .0019 | .2644 .0005 | .2260 .0017 | .2577 .0010 | .1572 .0018 | .1846 .0020 | .2114 .0014 | .2639 .0009 | .2866 .0012 | |
| MRR@10 | .2474 .0018 | .2448 .0020 | .2425 .0004 | .2038 .0015 | .2306 .0010 | .1353 .0021 | .1654 .0023 | .1875 .0015 | .2427 .0008 | .2646 .0015 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Value |
|---|---|
| Backbone layout | self-attn macro block mid block self-attn micro block |
| Routing granularity | Macro/mid group gate: session-wise; conditional expert gate: position-wise; micro: both position-wise |
| Expert groups | (Tempo, Focus, Memory, Popularity) |
| Experts per group | selected per dataset from |
| Expert activation | Top- groups; top- experts per selected group (6 active) |
| Cue dimensionality | 16 scalars per scope; 4 families 4 summaries each |
| Family | Macro (cross-session, H=5) | Mid (active session) | Micro (recent items) |
|---|---|---|---|
| Tempo | valid-history ratio inter-session gap mean pace pace trend | valid active prefix mean interval interval std. session age | valid local suffix last gap mean local gap gap shift vs. mid |
| Focus | theme entropy top-theme mass theme repeat theme shift | category entropy top-category mass switch rate category uniq. | current switch last-category mismatch suffix cat. entropy suffix cat. uniq. |
| Memory | repeat intensity adj.-category overlap adj.-item overlap repeat trend | item uniqueness repeat rate novel-item rate longest item run | last reconsumed suffix reconsume rate suffix item uniq. suffix longest run |
| Popularity | mean popularity pop. variability pop. entropy pop. trend | mean popularity pop. variability pop. entropy pop. trend | last-item popularity suffix pop. variability suffix pop. entropy pop. shift vs. mid |
| Dataset | Domain | Inter. (M) | Sess. (K) | Items (K) | Avg. len. |
|---|---|---|---|---|---|
| KuaiRec | Video | 3.44 | 183.3 | 6.5 | 18.8 |
| LastFM | Music | 14.23 | 658.6 | 426.2 | 21.6 |
| Retail Rocket | Browse | 0.37 | 41.8 | 26.6 | 8.7 |
| ML-1M | Movie | 0.57 | 14.2 | 3.1 | 40.4 |
| Foursquare | POI | 0.05 | 4.6 | 3.6 | 9.9 |
| Beauty | E-com. | 0.02 | 2.5 | 1.9 | 7.5 |
| Dataset | Model | Exec./ logical | Train examples/s | Full-sort ms/batch | Peak memory train / infer |
|---|---|---|---|---|---|
| Foursquare | SASRec-wide | 1.00 | 4,913.4 | 2.887 | 363.1 / 163.3 |
| Foursquare | Hidden MoE | 1.62 | 3,220.7 | 16.259 | 593.7 / 198.0 |
| Foursquare | RouteRec | 1.62 | 2,978.5 | 17.491 | 593.7 / 197.4 |
| KuaiRec | SASRec-wide | 1.00 | 5,121.5 | 4.722 | 652.7 / 288.3 |
| KuaiRec | Hidden MoE | 1.98 | 2,975.6 | 19.150 | 1,318.5 / 356.9 |
| KuaiRec | RouteRec | 1.98 | 2,841.4 | 20.289 | 1,319.4 / 356.0 |
| Parameter | Candidate sets |
|---|---|
| Shared optimization | |
| Learning rate | Log-uniform: KuaiRec ; Foursquare/Retail Rocket/ML-1M ; LastFM ; Beauty |
| Weight decay | |
| History length | or |
| Hidden / inner width | ; inner (or each) |
| Layers / heads | ; |