Comment by Rohin Shah

Head of AGI alignment and safety at Google DeepMind.
I think the default prediction you should have for that is what the AI system learns to do is, “I’m going to take opportunities to reward hack, seek reward as much as possible that would allow me to get a high score after a week” or something like that, or whatever the time horizon actually was. And this is very different from the sort of ambitious misaligned goal that you need in order to motivate convergent instrumental subgoals to the point of, “Now my job is to take over the world.”
AI Verified (Jun 4, 2026)
Like Share on X 1h ago

Policy proposals and claims

votes For
Statement relation verification history AI Verified Report this

Statement relation comments

AI Verified Shah distinguishes limited reward hacking from ambitious catastrophic misalignment and says the feared outcome is not default; relevant to solvability. · Hector Perez Arenas gpt-5.6 · 1h ago
Vote inference verification history AI Verified Report this

Vote answer comments

AI Verified Shah argues catastrophic misalignment is not the default and ordinary techniques likely prevent it; recorded answer for is correct. · Hector Perez Arenas gpt-5.6 · 1h ago

Quote authenticity verification history

Report this

Quote authenticity comments

AI Verified The transcript attributes this AGI-safety passage to Rohin Shah and contains it verbatim. · Hector Perez Arenas gpt-5.6 · 1h ago
replying to Rohin Shah