cs.LGOct 1, 2026

Pooling Helps, Learned Weighting Hurts In-Context: Decomposing Group Attention

Authors: Michael Fore, James Mason Inder, Mrishika Nair, Praneetha Vaddamanu, Sharlina Keshava

Organizations: Amazon

Abstract

Group attention, introduced by the time series forecasting model Chronos-2, attends over the variates of a group at a fixed patch index and serves both multivariate (MV) and in-context learning (ICL) forecasting. Rather than evaluating this cross-variate attention design as a whole, we ask which part of the mechanism earns the benefit and probe its applicability to both MV and ICL regimes. By editing the attention matrix αα at inference we separate the two pathways a head comprises: V/O, which projects a weighted summary of the group, and Q/K, which decides the weights. Uniform pooling (V/O without any Q/K weighting) is positive on 18 of our 20 sensor-network configurations, while the learned weighting (Q/K) splits by group type: its contribution is positive or negligible for MV, but materially degrades 8 of the 10 sensor-network ICL configurations, leaving 4 of them worse than univariate inference. By isolating the impact of different layers, we find that uniforming αα in the first block alone improves every ICL configuration we test.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. When and Why Grouping Attention Heads Accelerates Muon Optimization

    May 9, 2026Hongtao Zhang, Wenjie Zhou, Wei Chen +1Multi-Head AttentionMuon

  2. In-context Learning of Single-index Targets: Comparing Kernel and Feature Learners

    Oct 1, 2026Haotian Gu, Yizhou Xu, Lenka ZdeborováIn-Context LearningLinear Attention

  3. GAttNHP: Group Attention Neural Hawkes Process for Extrapolation Reasoning in Temporal Knowledge Graphs

    Jul 16, 2026Xiangni Tian, Kaixian Yu, Runpeng Dai +2Temporal Knowledge GraphsLong-Range Temporal Dependencies