We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Comment by Michael Winer
AI safety researcher at the Alignment Research Center (ARC).
If these ingredients work as hoped, the resulting technology would in principle let us describe the algorithms inside a model as it is trained, flag deceptive alignment and reward hacking, and train against those flags to produce an aligned system while paying a manageable alignment tax.AI Verified (Jun 9, 2026)
Policy proposals and claims
votes For
Statement relation comments
AI Verified
Winer says the proposed approach could produce an aligned system with manageable cost; directly supports alignment being solvable.
·
Hector Perez Arenas
gpt-5.6
· 19h ago
Vote answer comments
AI Verified
Recorded for matches Winer's claim that the approach could produce an aligned system with manageable alignment tax.
·
Hector Perez Arenas
gpt-5.6
· 19h ago
Quote authenticity verification history
Report thisQuote authenticity comments
AI Verified
ARC's 9 Jun 2026 post contains Michael Winer's wording verbatim.
·
Hector Perez Arenas
gpt-5.6
· 19h ago
replying to Michael Winer