How you can be reasonable at every step and still end up somewhere wrong
Picture a very slow game of telephone, except instead of whispering a sentence, each person makes one small, sensible-looking interpretation of the instruction before passing it on. Person one hears 'write a friendly product description' and passes on 'write a warm, casual product description.' Person two hears that and passes on 'write a playful, informal product description.' Person three passes on 'write a joke-filled product description.' Nobody in that chain did anything obviously wrong. Each step was a small, defensible nudge from the one before it. But five steps later, 'friendly' has quietly turned into 'joke-filled,' which was never actually asked for.
This is goal drift, and it's the hardest problem in this course because there's no single moment you can point to and say 'that's where it went wrong.' In a long-horizon task, an agent makes dozens of small decisions over many steps: how to interpret an ambiguous instruction, how to handle a minor course correction, how to fill in a detail the original goal didn't specify. Each one, judged only against the step right before it, looks perfectly reasonable. The problem is that small interpretation shifts compound. Enough small, individually reasonable nudges, stacked over enough steps, add up to a result that's meaningfully different from the original goal, even though nobody can point to the one decision that broke it.
