Comment by Vincent Cheng

AI safety researcher at METR and Cornell University student.
Auditing teams get access to a new model before internal deployment (and even automated audits for checkpoints during training), do a more extensive job, can catch egregious misalignment examples like the OAI/HF hack, and have stronger guarantees for what incidents they expect to appear in the wild.
Unverified (Sep 7, 2026)
Like Share on X 1h ago

Policy proposals and claims

votes For
Statement relation verification history Unverified Report this
No statement relation verification comments yet.
Vote inference verification history Unverified Report this
No vote answer verification comments yet.
replying to Vincent Cheng

Other authors to follow