TL;DR
Attention circuits in language models are not uniformly developed, with distinct emergence patterns observed across different architectures. The study utilized a participation-ratio spectral signal and capability-specific selectivity screens to analyze three 1B-class models.
✦ Why It Matters
Engineers can optimize training strategies by understanding the distinct emergence patterns of attention circuits in language models.
Key Takeaways
How It Works
The study employs a participation-ratio (PR) spectral signal to analyze attention-head circuits, focusing on their emergence during training. By applying a capability-specific selectivity screen, the researchers can identify different types of attention heads, such as induction and BOS-attractor heads, and track their development across various training checkpoints.
Related