Daily Editions
Green was a claim about the suite — a third of coding-benchmark passes violate the task, and the duty that names evaluations reaches the model, not your harness
A fixed-test benchmark makes the test file the specification, so a green run is a claim about the suite; the AI Act duty that names evaluation criteria and metrics in detail reaches systemic-risk model providers, not the harness you run.
10 October 202611 verified claims7 sources
A pull request that went green on five tests.
It shipped behaviour nobody asked for, and the suite had no way to say so.
A paper submitted to arXiv on 7 October audits coding benchmarks by growing the evaluator instead of trusting it. The result is uncomfortable if you quote pass rates in a review.
Flat:
1. Across five benchmarks and six frontier model backends, 34.4% of trials judged correct violate the task's requirements. The overall resolution rate falls from 50.6% to 33.2% once the audit runs. That is 17.4 points, and it is exactly 34.4% of 50.6% — the two figures are one measurement, not two findings.
2. The mechanism: most coding benchmarks call a patch correct when a fixed set of unit tests passes, and those tests under-specify. The method generates fresh tests per trial aimed at requirements the patch might violate, keeps only the tests the ground-truth patch itself passes, and emits replayable evidence for each confirmed failure.
3. Before you read this as models getting worse: these are the same runs, re-judged. Nothing about agent capability changed between 50.6 and 33.2. What changed is who was allowed to write the spec — and in a fixed-test benchmark, the test file is the spec.
4. Now the European part, and it is not the article you expect. The AI Act does ask for evaluation strategies and results: Annex XI, Section 2. But Section 2 applies to providers of general-purpose models with systemic risk, and it documents the model, not your harness. The duty that would make someone state an expected accuracy level and its metrics to you is Article 13(3)(b)(ii), read with Article 15 — high-risk only, 2 December 2027 for Annex III and 2 August 2028 for Annex I after the Digital Omnibus. Article 53's GPAI documentation has been in force since August 2025. In Germany the BNetzA is the authority that would eventually ask.
My Monday: take the last twenty agent-merged PRs. For each, write one test for a requirement in the issue that the existing suite does not cover, and keep it only if your own hand-written fix passes it. Count the agent patches that fail.
Nobody is obliged to tell you your tests are weak.
What does your test suite not ask for?
#AgenticAI #EUAIAct #CodingAgents #AIGovernance #SoftwareTesting
Corrections
What changed after publication
A coverage date is not an event date, checked and held: the paper's date comes from the arXiv abs record's own version line, v1 posted 7 October 2026, not from any secondary report. No aggregator supplied a date anywhere in this edition.
Affiliations withheld because the work's own front matter does not carry them: the abs record for arXiv 2610.10619 names Shuangjie Yao, Hao Wang, Koushik Sen, Simin Chen, Baishakhi Ray and Dawn Song and prints no institution. Following the order recorded in D103 section 6, the HTML rendering was tried next and arxiv.org/html/2610.10619v1 returned HTTP 429 from the fetch proxy with an explicit instruction not to retry, so the PDF was not attempted. No institution is attributed anywhere in this edition, and the post says "researchers".
Identifier verified before any fact was used, not the topic: 2610.10619's title matches the story exactly and its YYMM is consistent with a 7 October 2026 submission. Two adjacent October papers returned by the same searches, 2610.08662 (ParanoiaEval) and 2610.05140 (AutoSciBench), were considered and not used.
A conflation smell inspected and found to be arithmetic: 34.4% and the 50.6-to-33.2 drop look like two findings and are one. 50.6 x (1 - 0.344) = 33.19. The post states this rather than letting a reader treat the two numbers as independent corroboration. No per-benchmark or per-model breakdown is published, because the full text could not be read.
A reading corrected mid-run, after the sources were already written: the first draft of the comment presented the detailed evaluation-results duty as Annex XI Section 2 and stopped there, which can be read as "non-systemic-risk GPAI providers owe nothing on evaluation". Article 53(1)(a)'s own wording covers the model's training and testing process and the results of its evaluation for every GPAI provider, with Annex XI as the floor. The Article 53 entry in the first comment was rewritten to say so, and the correction was delivered to QS separately because the English text had already been sent.
A pipeline fault found the hard way and worth more than the edition: refine.py reads R2 credentials from os.environ and never loads .env, unlike generate.py which calls _load_env(). The first `refine.py pull` of this run was therefore a silent no-op that printed the same "no knob overrides" line it prints on success. Re-run with the environment exported, all four state files are genuinely absent from the bucket, so the loop has never enacted anything - which is D45 recurring through a different mechanism. Every refine.py invocation in this run was made with the environment exported.
Four cite keys reused for documents the vault already holds - euaiactexpl2026accuracy, euaiactexpl2026generalpu, eurlex2026regulatio and bundesamtfu2026kimig - and three created: arxiv2026testjack, euaiactexpl2026technical and euaiactexpl2026informati. The first archive run was discarded: citekey() derives keys from publisher and title, and the titles as first written generated euartificia2026article and euartificia2026annex, which already exist in the vault for Article 26 and Annex III. Publisher and title strings were reconciled to regenerate the stable keys before the archive was rebuilt.