Comment by Evan Hubinger

Research scientist at Anthropic working on model organisms of misalignment; former Research Fellow at MIRI.
Alignment auditing is starting to get really hard and we're going to need new techniques (e.g. interpretability-based) if we want to keep up.
AI Verified (Sep 1, 2026)
Like Share on X 20d ago

Policy proposals and claims

votes For
Statement relation verification history AI Verified Report this

Statement relation comments

AI Verified Quote says alignment auditing needs new interpretability-based techniques to keep up, strongly supporting interpretability requirements. · Hector Perez Arenas gpt-5.6 · 18d ago
Vote inference verification history AI Verified Report this

Vote answer comments

AI Verified Hubinger says interpretability-based techniques are needed to keep up with alignment auditing; recorded for matches. · Hector Perez Arenas gpt-5.6 · 18d ago

Quote authenticity verification history

Report this

Quote authenticity comments

AI Verified Verified verbatim in Axios; attributed to Evan Hubinger. · Hector Perez Arenas gpt-5 · 19d ago
replying to Evan Hubinger

Other opinions from Evan Hubinger

See all
AI alignment is solvable
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and ar...
more Unverifiable source (Sep 9, 2026)
Like Comment Share on X added 5d ago

Other authors to follow