A contact created, an email sent, and the deal stage exactly where it was on Friday. That is what a failure looks like when the grading is end-state.

On Wednesday Google shipped Gemini 4 Argon — to trusted cyber defenders first, through its Fairwind Program — and reported first place on AutomationBench at 51.3%.

Flat:

1. The headline numbers are Google's own: 77.9% on DeepSWE v1.1, a tie for first on CWE-bench v1 at 68%, and #1 on AutomationBench at 51.3%.

2. Here is the number that belongs beside it. Zapier's public AutomationBench leaderboard — 600 tasks, strict pass rate, 1.0 only if every assertion passes — tops out at Claude Opus 4.8 on 30.33%. When the benchmark was introduced in April, its paper put the best frontier models below 10%.

3. Before you conclude frontier agents now finish half your business workflows: there are two AutomationBenches. Zapier's grades binary, end-state, no partial credit. Artificial Analysis's AutomationBench-AA grades the average share of each task's objectives completed without triggering guardrail violations, and its top three sit at 71.3, 69.5 and 68.9. A 51.3% can only be first place on the strict metric — and Argon appears on neither public board. The claim may be perfectly good. Today it is Google's word, on Google's run.

4. And if your instinct is that a rule catches the gap: the duty that would ask you to supervise an agent is not live. Annex III high-risk — where the HR workflows in this benchmark would sit — moved to 2 December 2027 under the Digital Omnibus, Regulation (EU) 2026/1744, and Article 14 human oversight went with it. What binds today is Article 50 transparency, in force since 2 August 2026, with the machine-readable marking limb 62 days out on 2 December. That one makes you declare the machine, not check it. The limb that bites on a wrong CRM record is older and has no deferral: GDPR Article 5(1)(d), accuracy.

My Monday: make every agent write replayable, so a rerun is safe rather than doubled. One write scope per agent, per system. Reconcile the end state, not the step log — the step log is what the agent believed happened.

A benchmark score is a claim about a metric, and about who ran it.

Which of your numbers has a scorekeeper who is not the vendor?

#AgenticAI #EUAIAct #GDPR #AIGovernance #DevSecOps