AI From Scratch/Phase 10/Lesson 16/~60 minutes

Differential Attention (V2)

BuildPython (stdlib)No prerequisites

Softmax attention spreads a small amount of probability over every non-matching token. Over 100k tokens that noise adds up and drowns the signal. Differential Transformer (Ye et al., ICLR 2025) fixes it by computing attention as the differ...

Loading lesson page...