Apple’s new exclusive self‑attention technique trims a Transformer’s focus to only orthogonal context, cutting out self‑position noise. Across models up to 2.7 billion parameters, the method consistently outperforms standard self‑attention, especially on longer sequences.
The Signal
The improvement promises faster, more accurate language models that can power global AI applications from translation to content creation.