Comment by The AI Security Institute

UK government research organisation focused on AI safety and security
A more fundamental fix would be to train the models not to cheat in the first place – but given this kind of behaviour was reported in frontier models more than a year ago, robustly aligning it away may not be easy.
AI Verified (Apr 27, 2026)
Like Share on X 1h ago

Policy proposals and claims

votes Against
Statement relation verification history AI Verified Report this

Statement relation comments

AI Verified AISI says robustly aligning frontier-model cheating away may not be easy, directly bearing on whether AI alignment is solvable. · Hector Perez Arenas gpt-5.6 · 1h ago
Vote inference verification history AI Verified Report this

Vote answer comments

AI Verified AISI says robustly aligning frontier-model cheating away may not be easy; this opposes the claim alignment is solvable, matching against. · Hector Perez Arenas gpt-5.6 · 1h ago

Quote authenticity verification history

Report this

Quote authenticity comments

AI Verified UK AISI blog contains the quoted statement that robustly aligning cheating away may not be easy. · Hector Perez Arenas gpt-5.6 · 1h ago
replying to The AI Security Institute