AI auditing

News, analysis & learning

Analysis Published

Why a benchmark score does not establish agent safety

Tasks, tools, and compute budgets shape every evaluation score. Understanding those choices is the first step toward judging what a result says about deployment.

Inspect log viewer showing evaluation results and individual samples
Inspect log viewer, as shown in the project documentation. An example interface, not a result produced by Audit Commons. Inspect contributors · MIT
Read article →

News & Releases

See all →
Orientation

What makes AI auditable?

An editorial introduction explaining how evaluation, monitoring, and auditing differ by the questions they ask and the evidence they require.

Selected Resources

Full catalog →