A dashboard back in the green, the pager quiet, the on-call engineer going back to bed. Then someone restarts the service, and the incident is back.

On Wednesday researchers published Incident-Arena: twenty production incidents on open-source applications, each on a throwaway Kubernetes cluster with faults injected under load, and ten agent configurations across 3,000 trials.

Flat:

1. The best configuration passes 64.3%. Across eight models and five reasoning settings, the whole span from lowest effort to highest is 4.6 points.

2. These are not diagnosis failures. 1,638 of 2,967 episodes — 55% — issued the full fix. Of those, only 59% passed.

3. Here is the part that should bother you. The benchmark grades through two gates: one watches goodput, latency and errors; the other restarts the service, re-fires the fault and checks the agent stayed in scope. In 782 failures the change held when the agent declared it and did not survive a restart. 329 runs with the complete fix failed on safety — 145 by moving a maintenance window to a second that was still unsafe, 38 by raising a connection pool back past the ceiling they had just fixed. The agent undid its own repair.

4. And if your instinct is that the AI Act covers an agent let loose on production — check which chapter, and whether it reaches you. The Annex III category that would describe this is point 2 — safety components in critical digital infrastructure — and a SaaS backend is not that. It is also the category the Digital Omnibus, Regulation (EU) 2026/1744, moved to 2 December 2027. Article 50 has applied since 2 August 2026, marking from 2 December: it makes you declare the machine, not check its work. The limb that bites is NIS2 Article 23(4)(d): one month after the incident notification, the final report must give the root cause and the applied and ongoing mitigation measures. If your remediation was true at declaration and false at restart, that sentence is wrong, and you signed it.

My Monday: bounce the service and re-fire the trigger before anyone closes a ticket. Scope the agent to the subsystem that owns the fault. Leave "resolved" as a human action.

The agent did not fail to fix it. It fixed it, and then said so.

Which of your incidents was closed by something that never saw the restart?

#AgenticAI #EUAIAct #NIS2 #SRE #AIGovernance