A diff in the right file. A summary that sounds certain. A test suite that never asked the question the ticket was about.
That patch gets approved, and the approval means nothing.
On 1 October, researchers at UC Berkeley and Virginia Tech asked when a weaker model can reliably decide whether a coding agent's patch solves its issue. 411 execution-labelled traces from three agents, 101 controlled cases, six reviewer models from Llama-3.1-8B to Qwen3-235B.
Flat:
1. Hand the reviewer official execution evidence and five of six improve on both measures; two classified all 122 held-out traces correctly. Hand them structured-but-unchecked evidence — selected, confident, nothing verified — and both numbers move the wrong way together: more defects caught, more good patches rejected.
2. Before you conclude the fix is a bigger reviewer: reviewer size is not a consistent predictor of quality. The ladder spans 8B to 235B and the ranking is not monotonic. The lever is groundability — whether an independent check signed the evidence — not scale.
3. Official tests don't exist in deployment, so the authors froze a cascade: static errors plus generated tests that first fail on the unpatched repo. On held-out sets, catch 0.76 against over-rejection 0.66 on GPT-5.4, catch 0.80 against 0.67 on Gemini. Ten to thirteen points of separation. That is not a gate.
4. Now the European part, and it is not the article you expect. The oversight duty everyone reaches for is AI Act Article 14 — understand the system's limits well enough to spot anomalies. It moved to 2 December 2027 with Annex III, under Regulation (EU) 2026/1744, the Digital Omnibus. The Cyber Resilience Act has an Article 14 too, and it has bound manufacturers since 11 September 2026: 24 hours to warn on an actively exploited vulnerability, 72 to notify, a final report 14 days after a corrective measure is available. That clock starts at your fix, not your incident.
My Monday: make the gate produce a test that fails on the unpatched repo and passes after. Freeze one evidence format per gate. Track over-rejection next to catch — a gate nobody trusts gets switched off. And record which check signed the approval.
Oversight is not someone senior reading more carefully. It is a question your tooling can answer.
Which of your agent's approvals rests on a check that could have said no?
#AgenticAI #EUAIAct #CyberResilienceAct #CodeReview #AIGovernance