Loading lesson page...
Evaluation: Benchmarks, Evals, LM Harness
BuildPythonNo prerequisitesGoodhart's Law: when a measure becomes a target, it ceases to be a good measure. Every frontier lab games benchmarks. MMLU scores go up while models still can't reliably count the number of R's in "strawberry." The only eval that matters i...