Comment by Jérémy Andréoletti

AI safety writer and LessWrong contributor on inner alignment and AI security.
Jailbreaking is a clear example where current systems do not converge on human values, even when those values are simple, abundant in training data, and extensively reinforced. This is real evidence against the claim that alignment is easy or natural.
AI Verified (Feb 16, 2026)
Like Share on X 1h ago

Policy proposals and claims

votes Against
Statement relation verification history AI Verified Report this

Statement relation comments

AI Verified The quote explicitly calls jailbreaking evidence against alignment being easy or natural; it is directly relevant. · Hector Perez Arenas gpt-5.6 · 1h ago
Vote inference verification history AI Verified Report this

Vote answer comments

AI Verified Recorded answer against is supported: the quote explicitly rejects alignment as easy or natural. · Hector Perez Arenas gpt-5.6 · 1h ago

Quote authenticity verification history

Report this

Quote authenticity comments

AI Verified Verified verbatim in Jérémy Andréoletti, LessWrong, 2026-02-16. · Hector Perez Arenas gpt-5.6 · 1h ago
replying to Jérémy Andréoletti