Performance and evaluation efficiency plot from the optstop repository
Figure from the optstop authors: performance versus evaluation efficiency under their experimental settings. Toby D. Pilditch / UK AI Security Institute · MIT
News

AISI introduces optstop for adaptive evaluation stopping

The open-source tool uses statistical stopping rules to allocate evaluation trials. Its August announcement describes how evaluators can inspect and validate those decisions.

Published Announced
EvaluationBenchmarksCompute efficiencyAISIInspect
Read article →
Inspect log viewer showing evaluation results and individual samples
Inspect log viewer, as shown in the project documentation. An example interface, not a result produced by Audit Commons. Inspect contributors · MIT
Analysis

Why a benchmark score does not establish agent safety

Tasks, tools, and compute budgets shape every evaluation score. Understanding those choices is the first step toward judging what a result says about deployment.

Published
EvaluationAgent safetyBenchmarksCompute budgetAudit evidence
Read article →
Inspect log viewer showing evaluation results and individual samples
Inspect log viewer, as shown in the project documentation. An example interface, not a result produced by Audit Commons. Inspect contributors · MIT
Guide

How to read an agent evaluation report

Five questions about tasks, budgets, uncertainty, failures, and claims, applied to a worked example.

Published
EvaluationAgent auditingBenchmarksEvidence assessment
Read article →
Orientation

What makes AI auditable?

An editorial introduction explaining how evaluation, monitoring, and auditing differ by the questions they ask and the evidence they require.

Published
EvaluationMonitoringAuditing
Read article →