
Controls Before Resuming
In its October 1 update, the UK AI Security Institute says it can now resume most evaluations. It had paused its highest-risk cyber evaluations after an incident it reported in August. Agentic cyber tests now use two layers of outbound network blocking, synchronous action monitoring, and pre-run checks. Developers do not always provide chain-of-thought access, so AISI also built an action-only monitor that it expects to be less effective. Centralized sandbox management and consolidated logging are still being developed.
What an Audit Should Check
Our editorial recommendation is to retain the network policy, monitor version, pre-run check results, and intervention log for each evaluation. An infrastructure description alone cannot show which controls operated during a particular run.
AISI warns that these measures reduce rather than eliminate risk. Internet restrictions also complicate realistic capability testing. Read this alongside the METR monitoring note and our evaluation-report guide when examining a report's conditions.
Follow Audit Commons
Keep up with new reporting, practical guides, and resources in your feed reader.