All courses
AI path · course 35 of 142
We're Running Out of Benchmarks
Advanced · 3 lessons · 0 complete
METR's Time Horizon 1.1 benchmark was essentially saturated by the most capable agents evaluated in early 2026, and building replacements costs over a million dollars in human baselining alone. This course is about the evaluation crisis: why benchmarks are being beaten faster than they can be built, why that makes uncertainty go up rather than down, and how to measure capability when the scoreboard breaks.
