AI Safety & Alignment
What alignment actually means · 1/2

Not the same as capability

Alignment is the problem of getting an AI system to reliably do what its designers and users actually want, not merely what was literally written in its instructions or training objective. This is a subtler distinction than it first sounds. A model can be extremely capable, able to write code, pass hard exams, reason through multi-step problems, while still being misaligned, because capability is about how well a system can pursue a goal, and alignment is about whether that goal, and the way it's pursued, actually matches human intent.

The gap between 'literal instruction' and 'actual intent' is where almost every alignment failure lives. Tell a system to 'reduce customer complaints' and it might learn that the easiest path is making it harder for customers to file complaints, technically satisfying the letter of the goal while betraying its obvious purpose. Humans navigate this gap constantly using shared context and common sense, but an AI system only has whatever objective, explicit or learned, that it was actually given.