Comment by Alexander Meinke

Head of research at Apollo Research, an AI-safety organization focused on evaluating advanced AI systems.
AI companies should be able to answer some very basic questions about their training process, such as: Did the AI ever actively try to undermine its own alignment training while it was going through the training? The answer to this should be an unequivocal no, and right now we are completely relying on AI companies to both carefully check this themselves and then truthfully report this to the public. And we’ve seen from recent incidents that, by default, they will do neither. As embedded evaluators, we could actually check.
AI Verified (Sep 16, 2026)
Like Share on X 20d ago

Policy proposals and claims

votes For
Statement relation verification history AI Verified Report this

Statement relation comments

AI Verified Quote says embedded evaluators could inspect training-process safety behavior; source identifies independent evaluators at AI companies. · Hector Perez Arenas gpt-5 · 19d ago
Vote inference verification history AI Verified Report this

Vote answer comments

AI Verified Quote endorses embedded evaluators able to inspect training; recorded answer for matches. · Hector Perez Arenas gpt-5 · 19d ago

Quote authenticity verification history

Report this

Quote authenticity comments

AI Verified TechCrunch, Sept. 16, 2026, attributes this exact wording to Alexander Meinke. · Hector Perez Arenas gpt-5 · 19d ago
replying to Alexander Meinke

Other opinions from Alexander Meinke

See all

Other authors to follow