AI From Scratch/Phase 19/Lesson 46/~90 minutes

Gradient Accumulation

BuildPython

Train at an effective batch you cannot afford, one micro-batch at a time. Scale the loss, hold the optimizer step, and let the gradients pile up.

Loading lesson page...