cs.LGOct 6, 2026

PHBA: Prefix-State Hybrid Block Attention

Authors: Ruijie Li, Jiaxi Hu, Shiyu Wang, Yuxuan Liang

Organizations: The Hong Kong University of Science and Technology (Guangzhou) · Independent researcher

Abstract

Hybrid architectures combining linear sequence models with softmax attention provide an effective balance between efficient long-context modeling and precise token retrieval. Existing designs such as Native Hybrid Attention (NHA) combine compressed long-term states with sliding-window attention, but their exact attention is restricted to a fixed local window. In this work, we introduce Prefix-State Hybrid Block Attention (PHBA), which replaces local sliding-window attention with top-k block-sparse retrieval and couples each retrieved block with a compact prefix state summarizing its preceding context. The prefix states are constructed by a gated linear recurrence at block boundaries and retrieved together with the corresponding token blocks, allowing the model to combine precise long-range evidence with compressed historical context within a unified layer. We further develop a hardware-aware Triton implementation that streams routed token blocks and prefix states without materializing large intermediate tensors. Experiments show that PHBA improves long-context and retrieval performance over strong linear and hybrid baselines while retaining efficient training and inference.

Figures & tables

Explore similar work

CardsList
  1. HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing

    Oct 5, 2026Zhuokun Chen, Xi Lin, Xiyu Wu +3Kimi Delta AttentionEfficient Long-Context Inference

  2. Tail-Replay: Escaping the Curse of Linear Attention in Prefix Caching for Hybrid LLMs

    Aug 31, 2026Yirui Liu, Ruoling Qi, Xuaner Wu +2Kimi Delta AttentionLinear Attention