All courses
AI path Β· course 47 of 54
Mechanistic Interpretability
Advanced Β· 6 lessons Β· 0 complete
A trained neural network is a black box even to the people who built it: millions of numbers that somehow produce coherent language or accurate predictions, with no obvious explanation of how. This course is about the emerging field of reverse-engineering that black box, treating a model less like an oracle to be trusted and more like unfamiliar compiled code to be decompiled, understood, and checked for what it's actually doing.
