When Can Attention Heads Be Statically Defined?

9月 25, 2026·
Weixian Waylon Li
Weixian Waylon Li
,
Yintao Tai
,
Marcio Fonseca
,
Shay B. Cohen
· 1 分钟阅读时长
摘要
Some attention heads learn similar patterns across inputs. Reusing these patterns could reduce training cost by avoiding repeated query-key score computation and softmax. Through controlled pretraining comparisons, we identify Selective Attention Freezing (SAF), which selects heads with low attention-pattern variance and replaces their attention weights with fitted post-softmax means halfway through training. We represent these fixed patterns with absolute-position and relative-distance preferences, reducing storage from quadratic to linear in sequence length. A fused kernel reconstructs the patterns and executes ordinary-attention and replaced heads together. At matched training-token budgets, replacing 25% of attention heads gives 1.056x faster post-replacement optimiser updates at 124M parameters and 4K context, with a 0.77% perplexity increase. At 1B and 8K context, post-replacement updates are 1.068x faster on four GPUs including communication, with a 0.51% perplexity increase. The resulting models also accelerate long-input finetuning and causal prefill. After associative-recall adaptation, the 124M model with 25% replacement generalises to more key-value pairs at a fixed length better than ordinary attention and two pruning controls.
类型
出版物
Preprint 2026

Citation

@misc{li2026attentionheadsstaticallydefined,
      title={When Can Attention Heads Be Statically Defined?}, 
      author={Weixian Waylon Li and Yintao Tai and Marcio Fonseca and Shay B. Cohen},
      year={2026},
      eprint={2609.34650},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2609.34650}, 
}