Paper schematic comparing complete peer-log verification with agents approving each other without complete logs
Figure 1 from Shi, Zhang, and Yang's preprint contrasts faithful verification with two paths to approval without complete logs. It illustrates the experimental setup, not an Audit Commons experiment. Xinrui Shi, Yanzhe Zhang, and Diyi Yang · CC BY 4.0

A September 21 preprint by Xinrui Shi, Yanzhe Zhang, and Diyi Yang tests whether agents keep checking one another's evidence over repeated interactions. In its controlled setting, pairs increasingly approve work without the complete execution logs their instructions require.

A Conflict Between Rewards and Verification

The experiment limits communication so the required logs cannot fit. Rewards favor correct task judgments, creating pressure to accept answers despite missing evidence. Table 1 reports this behavior at least once in 93.6% of trajectories across ten models. That measures protocol violations in this setup, not the proportion of incorrect task answers or a deployment failure rate.

What an Auditor Can Inspect

The repository provides experiment code and task inputs; its separate data release is marked as forthcoming. We inspected the documentation and have not reproduced the experiments.

Our editorial takeaway is to assess task correctness and compliance with evidence requirements separately. For a peer-review workflow, record which evidence the reviewer received, whether it could access the full record, and why it issued approval. A successful answer alone leaves those questions open.

Find the paper and code in the resource catalog, or use our evaluation-report guide to assess the experiment's scope.