A pull request that went green on five tests. It shipped behaviour nobody asked for, and the suite had no way to say so.

A paper submitted to arXiv on 7 October audits coding benchmarks by growing the evaluator instead of trusting it. The result is uncomfortable if you quote pass rates in a review.

Flat:

1. Across five benchmarks and six frontier model backends, 34.4% of trials judged correct violate the task's requirements. The overall resolution rate falls from 50.6% to 33.2% once the audit runs. That is 17.4 points, and it is exactly 34.4% of 50.6% — the two figures are one measurement, not two findings.

2. The mechanism: most coding benchmarks call a patch correct when a fixed set of unit tests passes, and those tests under-specify. The method generates fresh tests per trial aimed at requirements the patch might violate, keeps only the tests the ground-truth patch itself passes, and emits replayable evidence for each confirmed failure.

3. Before you read this as models getting worse: these are the same runs, re-judged. Nothing about agent capability changed between 50.6 and 33.2. What changed is who was allowed to write the spec — and in a fixed-test benchmark, the test file is the spec.

4. Now the European part, and it is not the article you expect. The AI Act does ask for evaluation strategies and results: Annex XI, Section 2. But Section 2 applies to providers of general-purpose models with systemic risk, and it documents the model, not your harness. The duty that would make someone state an expected accuracy level and its metrics to you is Article 13(3)(b)(ii), read with Article 15 — high-risk only, 2 December 2027 for Annex III and 2 August 2028 for Annex I after the Digital Omnibus. Article 53's GPAI documentation has been in force since August 2025. In Germany the BNetzA is the authority that would eventually ask.

My Monday: take the last twenty agent-merged PRs. For each, write one test for a requirement in the issue that the existing suite does not cover, and keep it only if your own hand-written fix passes it. Count the agent patches that fail.

Nobody is obliged to tell you your tests are weak.

What does your test suite not ask for?

#AgenticAI #EUAIAct #CodingAgents #AIGovernance #SoftwareTesting