AI From Scratch/Phase 09/Lesson 06/~75 minutes

Policy Gradient — REINFORCE from Scratch

BuildPython

Stop estimating value. Parameterize the policy directly, compute the gradient of expected return, step uphill. Williams (1992) wrote it in one theorem. It is why PPO, GRPO, and every LLM RL loop exist.

Loading lesson page...