cs.LGOct 7, 2026

Evaluating Trajectory Features for Routing Final-Layer Attention

Authors: Yupeng Yao

Abstract

Attention routing requires a signal that predicts the value of attention on the current prefix. We evaluate whether hidden-state extrapolation error, curvature and error change improve this prediction beyond uncertainty, one-step displacement, position and state projections. Paired executions of the final attention layer supply signed next-token loss differences in frozen SmolLM3-3B-Base and Qwen3.5-4B-Base checkpoints. Utility-supervised routers are tested on 100 held-out PG-19 books at an identical causal 20 percent invocation quota. None of six prespecified comparisons shows a positive gain after familywise correction. In Qwen3.5, a parameter-matched fixed-projection control lowers NLL by 0.00356 nats/token relative to the trajectory router (95 percent interval 0.00218 to 0.00487). Secondary results depend on the operation removed, feature location and scoring horizon; frozen thresholds also drift substantially at longer horizons. Actual selected-query execution yields small long-sequence latency reductions with increased NLL, while learned routers remain slower during cached continuation. The study identifies limits on the incremental value of these trajectory summaries and separates allocation quality from measured inference benefit.

Figures & tables

Appendix figures & tables36 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation

    Jul 7, 2026Andrii Balashov, Olena PonomarovaKey-Value Cache QuantizationMixture-Of-Experts

  2. The Routing Plateau: Understanding the Accuracy Limits of LLM Routers

    May 27, 2026Yifan Lu, Qiyue Zhang, Shenrun Zhang +4Large Language Model RoutingBottlenecks