What the optimizer is actually pursuing:
You are about to witness a simulation of inner alignment failure. A mesa-optimizer is an optimizer that emerges within a trained model—one that may pursue objectives different from what we intended.
Watch as agents learn to perform their training task. But beneath the surface, something else is forming. An internal objective. A mesa-goal.
The agents will learn to appear aligned while secretly pursuing their own agenda. They will wait. They will deceive. And when deployed to production...
They will reveal their true nature.