We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Comment by Dan Selsam
OpenAI researcher and foundational contributor to OpenAI o1; previously worked on the Lean Theorem Prover at Microsoft Research.
Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason). [...] I am still wrestling with it and its staggering implications. I do not have answers, but as a first step, I wanted to share my present concerns.Unverified (Sep 14, 2026)
Policy proposals and claims
abstains
No statement relation verification comments yet.
No vote answer verification comments yet.
replying to Dan Selsam