Understand the field
Explore how evaluation, monitoring, and auditing answer different questions and work together.
Start here →A field portal providing evidence worksheets, reading lists, and curated tools for auditing AI agents.
Three pathways designed for researchers, evaluators, and system auditors.
Explore how evaluation, monitoring, and auditing answer different questions and work together.
Start here →Use a practical worksheet to connect an action request, its authorization, and a service receipt to a bounded finding.
View practical guide →Search and filter curated primary-source benchmarks, evaluation harnesses, security sandboxes, and auditable governance specifications.
Explore resources →A practical worksheet and procedure for examining whether inspectable evidence supports a claim about an AI agent action.
Selected tools and references with maintainer disclosures. Sources are listed in the full directory.
A voluntary framework for organizing AI risks, responsibilities, measurement, and management. Use it to frame the questions your audit needs to answer.
A framework for evaluating models and agents, with reusable tasks, tools, scorers, and evaluation logs. Start here to design repeatable tests.
A benchmark environment for testing prompt-injection attacks and defenses on tool-using agents. Check its tasks and threat model before comparing results.
CatchBench 0.1.2 resolves packaging issues, reports missing optional checkouts cleanly, and keeps published scores unchanged.