A voluntary framework for organizing AI risks, responsibilities, measurement, and management. Use it to frame the questions your audit needs to answer.
Auditing Resources & Toolchains
Curated benchmarks, risk frameworks, evaluation toolchains, and primary reference materials for AI auditing.
No resources match your search or filter criteria.
A framework for evaluating models and agents, with reusable tasks, tools, scorers, and evaluation logs. Start here to design repeatable tests.
A benchmark environment for testing prompt-injection attacks and defenses on tool-using agents. Check its tasks and threat model before comparing results.
A maintained reading list of papers, tools, benchmarks, and standards covering agent reliability, monitoring, failure attribution, and accountability.
A benchmark for auditing agent failures. The offline PRE quickstart runs without model API keys; other benchmark paths have additional setup requirements.
An ecosystem page connecting tools for inspecting, enforcing policies on, and reviewing agent actions. Follow the component repositories for current capabilities.