We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Comment by Stephen Casper
AI safety researcher at the UK AI Security Institute, researching safeguards for open-weight AI models.
Third-party evaluations for frontier AI have mostly tested models through external interfaces before deployment. But the risks from frontier AI models depend on how their developers use and govern them internally. Recently, CEOs of frontier AI companies have committed to hosting embedded assessments. These assessments would give independent evaluators employee-like access to a developer's internal systems, staff, and documentation. [...] We recommend that frontier AI developers begin hosting embedded assessments now, covering at least three areas central to managing risks from internal AI use: internal agent monitoring, internal agent security controls and permissions, and model alignment. [...] Evaluators should have parity of access with internal staff conducting similar risk assessments or senior alignment researchers, as appropriate.Unverified (Sep 21, 2026)
Policy proposals and claims
votes For
No statement relation verification comments yet.
No vote answer verification comments yet.
replying to Stephen Casper