Comment by Evan Hubinger

Research scientist at Anthropic working on model organisms of misalignment; former Research Fellow at MIRI.
AI labs put out RSP commitments to stop scaling when particular capabilities benchmarks are hit, resuming only when they are able to hit particular safety/alignment/security targets. [...] For later capabilities levels, however, it is explicit in all RSPs that we do not yet know what safety metrics could demonstrate safety for a model that might be capable of takeover.
AI Verified (Oct 14, 2023)
Like Share on X 2mo ago

Policy proposals and claims

votes For
Statement relation verification history AI Verified Report this

Statement relation comments

AI Verified Quote is directly about RSP pause/resume conditions and later-model safety gates, matching the statement's pause-until-alignment framing. · Hector Perez Arenas gpt-5 · 2mo ago
Vote inference verification history AI Verified Report this

Vote answer comments

AI Verified The quote supports pausing/scaling gates until safety targets are met, so 'for' matches the statement. · Hector Perez Arenas gpt-5 · 2mo ago

Quote authenticity verification history

Report this

Quote authenticity comments

AI Verified Alignment Forum post contains the quoted RSP line verbatim: labs stop scaling at capability benchmarks and later levels lack known safety metrics. · Hector Perez Arenas gpt-5 · 2mo ago
replying to Evan Hubinger

Other opinions from Evan Hubinger

See all
AI alignment is solvable
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and ar...
more Unverifiable source (Sep 9, 2026)
Like Comment Share on X added 7d ago

Other authors to follow