Loading lesson page...
Reflexion: Verbal Reinforcement Learning
BuildPython (stdlib)Gradient-based RL needs thousands of trials and a GPU cluster to fix a failure mode. Reflexion (Shinn et al., NeurIPS 2023) does it in natural language: after each failed trial, the agent writes a reflection, stores it in episodic memory,...