Daily Editions
The grader travelled — RL on 1,700 coding tasks transfers across harnesses, and the instrument with no compute threshold lands on 9 December
A rank-32 LoRA is far below the AI Act's downstream-modifier compute threshold, so the Act leaves a team that post-trains an open-weight coding model alone. Directive (EU) 2024/2853 has no threshold: from 9 December 2026 software is a product and substantial modification makes the modifier the manufacturer. The transferable artefact in the paper is the regression-zeroing grader, not the weights.
4 October 202613 verified claims8 sources
A pull request that does every single thing the ticket asked, and quietly breaks a test nobody reran.
That is the failure two years of agent work has not fixed.
On Thursday, researchers published a post-training run aimed straight at it: reinforcement learning alone, on an open-weight model, checked on benchmarks released after the training data was collected.
Flat:
1. One epoch of RL on a rank-32 LoRA adapter over 1,700 tasks. Terminal-Bench 2.1 went 67.4 to 82.0. SWE-Bench Pro went 60.1 to 64.8. Same method, same run — fifteen points in one place, four in another.
2. The base is Kimi K2.7 Code: 1T parameters, 32B active, open weights under a Modified MIT licence. Something you can host.
3. The part worth copying is the reward, not the weights. Reward is the fraction of target checks passed, and it drops to zero if any pass-to-pass test fails. Not partial credit — zero. The gains held on two harnesses never used in training, and on the three benchmark sets released after the training data was collected (p = 0.004). Median trajectories got 24-35% shorter in agent steps. It did not get cleverer. It stopped doing the extra thing.
4. Now the European part, and it is not the chapter you expect. A rank-32 adapter is orders of magnitude below the compute threshold that would make you the provider of a general-purpose AI model, so on the AI Act, tuning an open-weight model and shipping it costs you nothing. Directive (EU) 2024/2853 has no threshold. From 9 December 2026 software is a product; Article 8(2) treats anyone who substantially modifies a product and then puts it on the market as its manufacturer; Article 7(2)(c) tells the court to weigh its ability to keep learning after release. Germany's transposing bill excludes open-source software supplied outside a commercial activity. That is the base model. It is not your build of it.
My Monday: make the acceptance gate return zero when one old test fails, not 0.9. Keep the pass-to-pass suite physically separate from the fail-to-pass one. Date-stamp every eval set against your training cut. And write the modification down, because from December that note is evidence.
The run did not teach a model. It taught a grader, and the grader is the part you can copy.
What does your acceptance gate do when one old test fails?
#AgenticAI #EUAIAct #OpenWeights #ProductLiability #AIGovernance
Corrections
What changed after publication
A primary full text that could not be read, recorded rather than papered over: the HTML full text of arXiv:2610.00890 at arxiv.org/html/2610.00890v1 returned HTTP 429 from the fetch proxy, which instructed that the page not be retried. Every figure in this edition therefore comes from the paper's own abstract record on arxiv.org/abs/2610.00890, which carries all six before-and-after pairs, both p-values, the reward definition, the task split and the trajectory-length reduction verbatim. No number rests on an unread page, and no figure from the body is used.
Affiliations withheld because the work's own front matter does not carry them: the arXiv abstract record names Sushant Mehta, Logan Ritchie and Edwin Chen and prints no institution. The HTML full text, which would be the next place to look, could not be retrieved. No institution is attributed anywhere in this edition and none was taken from an aggregator.
Identifier verified before any figure was used, not the topic: arXiv 2610.00890's title matches the story and its YYMM is consistent with a 1 October 2026 submission. The adjacent benchmark papers returned by the same keyword search were not used, and no figure in this edition comes from any other record.
A figure pair reported but deliberately not led on: two of the six gains start from a floor, Terminal-Bench 3 at 1.4 to 12.1 and Terminal-Bench 4 at 0.0 to 7.6. A jump off zero is not the same kind of fact as 67.4 to 82.0. The post quotes neither; both are stated in the first comment with the reason, and the argument rests on the paper's pooled significance instead.
An extraction error caught against a second rendering of the same primary: a first read of the EUR-Lex ELI page for Directive (EU) 2024/2853 reported the transposition deadline as 8 December 2024. That is the Directive's entry-into-force date appearing in page metadata, not the Article 22 deadline. The CELEX HTML rendering gives 9 December 2026, which matches the vault's existing verified claim from the 21 August edition. The extraction was the error, not the source.
A cite-key collision caught before archiving, not after: the Bundestag Drucksache 21/4297 was first titled so that it derived deutscherbu2026produktha, a key the vault already holds for a different Deutscher Bundestag page, the heute-im-bundestag note on the same bill. Both of this edition's Bundestag sources would have collapsed onto one key. The Drucksache was retitled to derive deutscherbu2026entwurf, which is the key the vault already holds for that exact PDF. Six keys reused for documents the vault already has and two created, arxiv2026crossbenc and huggingface2026kimikcode.
A settled position in the vault checked before arguing against it: an early draft of this edition argued that an autonomous coding agent generating code falls under the Article 50(2) machine-readable marking duty landing on 2 December 2026. The 30 July edition established that the Commission's Article 50 guidelines of 20 July 2026 carve source code out of that duty. The argument was dropped rather than published, and Article 50 appears here only as the date standing next to the Directive's.