Illustrated isolated evaluation server, two network barriers, tool-action inspection, and a pre-run checklist
AI-generated editorial illustration of layered network blocking, tool-action monitoring, and pre-run checks described by AISI. The scene is conceptual; it does not depict actual infrastructure or measure control effectiveness. Audit Commons (AI-generated editorial illustration) · Original AI-generated editorial illustration; reuse license pending

Controls Before Resuming

In its October 1 update, the UK AI Security Institute says it can now resume most evaluations. It had paused its highest-risk cyber evaluations after an incident it reported in August. Agentic cyber tests now use two layers of outbound network blocking, synchronous action monitoring, and pre-run checks. Developers do not always provide chain-of-thought access, so AISI also built an action-only monitor that it expects to be less effective. Centralized sandbox management and consolidated logging are still being developed.

What an Audit Should Check

Our editorial recommendation is to retain the network policy, monitor version, pre-run check results, and intervention log for each evaluation. An infrastructure description alone cannot show which controls operated during a particular run.

AISI warns that these measures reduce rather than eliminate risk. Internet restrictions also complicate realistic capability testing. Read this alongside the METR monitoring note and our evaluation-report guide when examining a report's conditions.