A pull request that does every single thing the ticket asked, and quietly breaks a test nobody reran. That is the failure two years of agent work has not fixed.

On Thursday, researchers published a post-training run aimed straight at it: reinforcement learning alone, on an open-weight model, checked on benchmarks released after the training data was collected.

Flat:

1. One epoch of RL on a rank-32 LoRA adapter over 1,700 tasks. Terminal-Bench 2.1 went 67.4 to 82.0. SWE-Bench Pro went 60.1 to 64.8. Same method, same run — fifteen points in one place, four in another.

2. The base is Kimi K2.7 Code: 1T parameters, 32B active, open weights under a Modified MIT licence. Something you can host.

3. The part worth copying is the reward, not the weights. Reward is the fraction of target checks passed, and it drops to zero if any pass-to-pass test fails. Not partial credit — zero. The gains held on two harnesses never used in training, and on the three benchmark sets released after the training data was collected (p = 0.004). Median trajectories got 24-35% shorter in agent steps. It did not get cleverer. It stopped doing the extra thing.

4. Now the European part, and it is not the chapter you expect. A rank-32 adapter is orders of magnitude below the compute threshold that would make you the provider of a general-purpose AI model, so on the AI Act, tuning an open-weight model and shipping it costs you nothing. Directive (EU) 2024/2853 has no threshold. From 9 December 2026 software is a product; Article 8(2) treats anyone who substantially modifies a product and then puts it on the market as its manufacturer; Article 7(2)(c) tells the court to weigh its ability to keep learning after release. Germany's transposing bill excludes open-source software supplied outside a commercial activity. That is the base model. It is not your build of it.

My Monday: make the acceptance gate return zero when one old test fails, not 0.9. Keep the pass-to-pass suite physically separate from the fail-to-pass one. Date-stamp every eval set against your training cut. And write the modification down, because from December that note is evidence.

The run did not teach a model. It taught a grader, and the grader is the part you can copy.

What does your acceptance gate do when one old test fails?

#AgenticAI #EUAIAct #OpenWeights #ProductLiability #AIGovernance