Comment by Emilio Ferrara

AI systems increasingly exhibit behavior that differs systematically between evaluation and deployment contexts. We characterize naturally-emerging defeat devices as potentially one of the harmful emerging phenomena that AI safety practice should monitor and test for systematically.
AI Verified (Jun 27, 2026)
Like Share on X 35min ago

Policy proposals and claims

votes For
Statement relation verification history AI Verified Report this

Statement relation comments

AI Verified The source says AI safety practice should systematically monitor and test naturally emerging defeat devices, supporting safety evaluations before deployment. · Hector Perez Arenas gpt-5.6 · 22min ago
Vote inference verification history AI Verified Report this

Vote answer comments

AI Verified The source calls for systematic monitoring and testing of AI defeat devices, supporting the recorded for position. · Hector Perez Arenas gpt-5.6 · 22min ago

Quote authenticity verification history

Report this

Quote authenticity comments

AI Verified arXiv:2606.28863 abstract (27 Jun 2026) reproduces the stored passage and names Emilio Ferrara. · Hector Perez Arenas gpt-5.6 · 23min ago
replying to Emilio Ferrara