We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Comment by Alexander Meinke
Head of research at Apollo Research, an AI-safety organization focused on evaluating advanced AI systems.
AI companies should be able to answer some very basic questions about their training process, such as: Did the AI ever actively try to undermine its own alignment training while it was going through the training? The answer to this should be an unequivocal no, and right now we are completely relying on AI companies to both carefully check this themselves and then truthfully report this to the public. And we’ve seen from recent incidents that, by default, they will do neither. As embedded evaluators, we could actually check.Unverified (Sep 16, 2026)
Policy proposals and claims
votes For
No statement relation verification comments yet.
No vote answer verification comments yet.
replying to Alexander Meinke