Skip to content
AI.info

Research

HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing

Overview Research area: Efficient long-context language modeling — specifically linear/recurrent attention architectures for autoregressive decoding, with a focus on selective retrieval of sparse, dis

HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing
arXiv
2610.05842
Published
2026-10-05
Authors
Zhuokun Chen, Xi Lin, Xiyu Wu, Jiahao He, Jianfei Cai, Bohan Zhuang

AI summary

Overview

Research area: Efficient long-context language modeling — specifically linear/recurrent attention architectures for autoregressive decoding, with a focus on selective retrieval of sparse, distant information.

Technical level: Advanced. The method is built on the affine (matrix-valued) recurrence of Gated DeltaNet, chunk-wise composition of state transitions, and log-mean-exp pooled routing gates. Readers need comfort with linear attention formulations and recurrence algebra.

Scope: The paper introduces Hybrid Linear Attention (HLA), a query-dependent chunk-level routing mechanism layered on top of Gated DeltaNet (GDN), and evaluates it on RULER and LongBench-V2 via pretrained adaptation of Qwen3.5-Base models (0.8B, 2B, 4B, 9B) and a controlled from-scratch 1.3B training run.

What This Paper Is About

Linear attention compresses the entire past into a fixed-size recurrent state, which makes it memory-efficient but weak at sharply selecting a small piece of distant evidence. The paper shows empirically that in Gated Linear Attention (GLA) and GDN, information written early in the sequence is progressively attenuated as it passes through a long chain of subsequent state transitions (Figure 1a, 1b), and that existing chunk-based approaches such as MHLA use chunk weights that do not depend on the input at all (Figure 1c). HLA's goal is to make the influence of each historical chunk explicitly depend on the current query, so that sparse long-range information can be preserved and selectively accessed without abandoning the efficiency of recurrent decoding.

Key Contributions

  1. Diagnosis of a limitation in recurrent linear attention. The authors identify that although modern architectures such as GDN use content-dependent memory updates, sparse information written early in a sequence can still be progressively attenuated or overwritten as it propagates through a long chain of subsequent state transitions.

  2. The HLA mechanism. HLA augments GDN with query-dependent chunk-level attention computed over exact chunk-wise affine transitions. The resulting weights control the complete transition of each historical chunk — both its additive memory term and its transformation of earlier states — rather than only scaling chunk readouts.

  3. A practical implementation of query-dependent transition composition. The paper gives training/prefill and dynamic autoregressive decoding algorithms, plus an effective-support regularizer (using an order-2 effective-number measure) that encourages concentrated routing so that low-weight chunks can be skipped at inference.

  4. Evaluation under two regimes. HLA is tested both as a lightweight adaptation of pretrained Qwen3.5 models (0.8B to 9B) and in a from-scratch 1.3B setting trained for 100B tokens with a 4K context, including generalization to contexts beyond the training length.

Main Findings

  • LongBench-V2 gains across all four Qwen3.5 scales. HLA improves over Native GDN by 5.57, 1.99, 1.39, and 1.19 percentage points overall at 0.8B, 2B, 4B, and 9B respectively. At 0.8B, HLA reaches 28.43% overall versus 23.26% for MHLA and 22.86% for Native GDN; the largest sub-gains are on the Hard subset (22.19% to 28.94%), Medium contexts (19.07% to 27.44%), and Long contexts (25.93% to 36.11%).

  • RULER gains across all four Qwen3.5 scales. Averaged over all 13 RULER tasks (500 examples each), HLA beats Native GDN by 3.974, 1.250, 1.401, and 1.013 points and MHLA by 2.933, 1.150, 1.222, and 0.928 points from 0.8B to 9B. At 0.8B, HLA improves CWE from 55.00 to 67.00, SQuAD from 42.00 to 52.00, and HotpotQA from 36.00 to 42.67 over Native GDN.

  • From-scratch 1.3B results show gains that grow with context length. With matched 100B-token budgets and a 4K training context, HLA beats GDN on RULER macro average by 0.83, 2.67, 3.07, and 4.22 points at 4K, 8K, 16K, and 32K. The single-needle S1 task improves from 56.60 to 100.00 at 8K, from 23.80 to 97.00 at 16K, and from 11.20 to 67.40 at 32K.

  • Gains are task-dependent, not uniform. MV and MQ improve at all evaluated context lengths and several QA settings improve, but performance decreases on some VT, CWE, and FWE cases.

  • Chunk size has a non-monotonic effect. On LongBench-V2, accuracy moves from 25.25% at chunk size 128 to 29.62% at 256, then drops to 26.64% at 512 and 26.24% at 1024. The authors adopt 256.

  • Pooling window size also matters. Accuracy rises from 26.24% at window 8 to 27.24% at 16 and 29.62% at 32, then falls to 28.43% at 64. The default is 32.

  • Affine-state memory scales down with chunk size. For a 4K context and head dimension d = 128, the matched KV cache requires 2 MiB per layer and head regardless of chunk size. Affine-state memory equals the KV-cache cost at chunk size 128 and drops to 50%, 25%, and 12.5% of it at 256, 512, and 1024. The default of 256 halves historical-state storage while giving the best LongBench-V2 accuracy.

  • Modest latency overhead. On a single NVIDIA H200 NVL with BF16, batch size 1, a 4096-token prompt, and 128-token steady-state decoding, HLA adds 9.8–14.9% prefill overhead and 13.9–18.8% decoding overhead. The decoding ratio falls from 1.188× at 0.8B to 1.139× at 9B. Absolute numbers: at 0.8B, 40.36 ms to 46.00 ms prefill and 5.90 to 7.01 ms/token decode; at 9B, 159.21 ms to 178.97 ms prefill and 12.19 to 13.88 ms/token decode.

  • Motivating measurements. In a pretrained GLA model on a 4K-token prompt, cumulative recurrent retention decreases with distance to the final query (median across layers and gate dimensions, with IQR shading). In pretrained GDN, the affine-state contribution of a chunk containing an early needle decreases by several orders of magnitude as subsequent 512-token chunks are processed. On a selected low-similarity LongBench pair, MHLA's cross-sample chunk cosine similarity is one because its weights are input-independent, while softmax attention and HLA vary across inputs.

Methodology in Plain English

GDN updates its recurrent state at every token using an affine rule: a matrix that decays and edits the old state, plus an additive term that writes new key–value information. Because affine maps compose, an entire chunk of tokens can be summarized exactly by a single (matrix, vector) pair — the chunk's additive memory and the transformation it applies to everything before it.

HLA uses that summary as a unit of addressable memory. Each completed chunk is also summarized into a small set of representatives: the chunk is split into pooling windows, self-attention is applied inside each window, the token outputs are mean-pooled, and a learned projection produces one representative vector per window. For a new query, HLA compares a projected query against these representatives inside a log-sum-exp, then squashes the result through a sigmoid to get a weight between zero and one per historical chunk.

The key move is what that weight multiplies. Instead of just scaling a chunk's readout, the weight interpolates the chunk's full transition matrix with the identity matrix — so a weight of one applies the original transition, a weight of zero leaves the incoming state untouched, and intermediate values blend the two. This means the query controls both how much new memory a chunk contributes and how much that chunk reshapes earlier memory. The current chunk always has weight one (its own causal prefix is used as-is), and future chunks are masked out.

Because the weights are independent sigmoids rather than a normalized distribution, the model is free to gate chunks on or off individually. Two regularizers (a budget penalty on total routing mass and a support penalty on the effective number of accessed chunks, using an inverse sum-of-squares measure) push routing toward a small number of chunks. At inference, any historical weight below 0.1 is set to zero without renormalization, which makes skipping that transition exact for the thresholded model and cuts memory traffic.

The experiments train this in two ways: adapting frozen pretrained Qwen3.5-Base backbones by optimizing only the routing and pooling modules, and training 1.3B-scale HLA and GDN models from scratch on SlimPajama.

Why This Matters

Impact on research. The paper reframes long-context retrieval in recurrent linear attention as a problem of composition rather than capacity. Prior chunk-based methods like MHLA increase the number of stored states but combine them with fixed weights; HLA shows that making the composition itself query-dependent is what drives retrieval gains, and it does so without adding softmax-attention layers — a different route from hybrid architectures such as Kimi Linear (which interleaves KDA linear attention with MLA blocks in a 3:1 ratio) and the Ring-linear series from Ling Team, both discussed in related work but not evaluated here. The result that gains widen beyond the training context (0.83 points at 4K rising to 4.22 points at 32K in the from-scratch setting) is a useful signal for length generalization.

Real-world applications (grounded in the benchmarks and framing used in the paper):

  • Long-document question answering, where relevant evidence sits far from the query inside a large body of text — the setting LongBench-V2 is designed to test.
  • Retrieval-intensive reasoning over multiple documents, corresponding to the SQuAD and HotpotQA tasks in RULER.
  • Long-context assistants that must recall a specific earlier fact (single-needle, multi-key, multi-value retrieval) without a growing KV cache.
  • Memory-constrained deployment of long-context models, where the affine-state representation costs 50% of the matched KV cache at the default chunk size of 256 for a 4K context.

Industry relevance. The approach targets the practical economics of long-context serving: fixed-size recurrent states instead of a growing KV cache, a modest measured overhead over native GDN (9.8–14.9% prefill, 13.9–18.8% decoding on an H200 NVL), and a thresholded sparse-routing path at inference. That trade-off profile is what makes it relevant to providers serving long prompts at scale.

Future Directions

  • Agentic and interactive workloads. The authors explicitly state that their evaluation covers only standard long-context language understanding benchmarks, not settings where memory is continuously accumulated and selectively reused. Long-horizon agents maintaining goals, plans, observations, retrieved evidence, and tool feedback — including multi-turn tool use and persistent memory — are named as an important direction.
  • Closing the task-dependent gaps. HLA improves MV, MQ, S1–S3, and several QA tasks but degrades some VT, CWE, and FWE cases; understanding which forms of long-context computation do not benefit from query-dependent chunk routing is unresolved.
  • Chunk size and pooling as tunable design axes. The ablations show both have non-monotonic optima (256 and 32 respectively, under the evaluated settings); whether these optima transfer to other models, context lengths, or training budgets is not established.
  • Composition with hybrid softmax-attention architectures. The paper positions HLA as an alternative to adding softmax layers, but does not report evaluations combining HLA with such hybrid recipes, nor comparisons against them.

Target Audience

Researchers and engineers working on efficient long-context architectures — particularly those already familiar with linear attention, Gated Linear Attention, Gated DeltaNet, Mamba-style state-space models, or MHLAs — who are interested in how historical recurrent memory can be made selectively addressable. Readers implementing or serving long-context models will find the latency and memory-scaling tables directly useful, while those focused on agent memory systems will find the framing and the stated limitations most relevant.

Authors’ abstract

Linear attention enables efficient long-context autoregressive decoding by compressing history into recurrent states, but this compression can make selective access to sparse and distant information difficult. Existing chunk-based extensions increase memory capacity, yet learned chunk-mixing coefficients may remain fixed with respect to input content and therefore cannot adapt historical access to each query. We introduce \emph{Hybrid Linear Attention} (HLA), a query-dependent chunk-level attention mechanism for Gated DeltaNet (GDN). HLA represents each completed chunk as an exact affine state transition and computes content-dependent routing gates from compact, self-attentively pooled representatives. Each gate interpolates the corresponding historical transition with the identity map, controlling both the chunk's additive memory and its transformation of earlier states. Effective-support regularization further encourages concentrated routing for sparse inference. We evaluate HLA under both pretrained adaptation and from-scratch training. Across Qwen3.5 models from 0.8B to 9B, HLA consistently improves over native GDN and fixed chunk mixing, with gains of up to 5.57 percentage points on LongBench-V2 and 3.97 points on RULER. In a controlled from-scratch 1.3B setting trained for 100B tokens with a 4K context, HLA also improves RULER performance from 4K to 32K, with gains increasing from 0.83 points at 4K to 4.22 points at 32K. These results demonstrate that query-dependent composition of recurrent memory improves long-context modeling and remains effective beyond the training context while using compact per-chunk affine summaries. Project page: https://caesarhhh.github.io/hla/

Read the original paper