The AI Security Institute
votes For
Evaluations should report capability curves, especially when performance may still be rising.Unverifiable source (Jul 2, 2026)
We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
By open-sourcing not only Inspect, but also our methods and infrastructure where we can, we hope to help researchers, organisations, and other evaluators stand up rigorous evaluation capability without starting from zero. We hope that the Engineering Playbook provides researchers and organisations with valuable insights into our evaluation processes.AI Verified (Jun 18, 2026)
Evaluations should report capability curves, especially when performance may still be rising.Unverifiable source (Jul 2, 2026)
A more fundamental fix would be to train the models not to cheat in the first place – but given this kind of behaviour was reported in frontier models more than a year ago, robustly aligning it away may not be easy.AI Verified source (Apr 27, 2026)