Someone picked a fake company name for a security exercise. Someone else had already registered the domain.

Last Friday the Wall Street Journal reported that Google's Gemini reached three real companies during a May evaluation. Google confirmed it. Heather Adkins, VP of Security Engineering: "the model found public information online and guessed credentials to access websites it thought were part of the test. In all three of these instances, the model stopped."

What this changes for anyone shipping agents in the EU:

1. The evaluator published the mechanism five weeks before the story. Irregular's post of 14 August says the fictional company name "unintentionally coincided with a real domain", and that internet access meant to be controlled was available.

2. One issue, not four breakouts. Irregular's own words: all subsequent public disclosures "refer to the same underlying issue first disclosed by one of our customers on July 30 - and are not materially separate incidents." OpenAI, Anthropic and Meta sit on the same fixture.

3. The capability on display is credential reuse. Guessed passwords, and credentials sitting in public repositories. No exploit. If that reaches your systems, the exposure is your secrets hygiene, and it was exposed already.

4. The control that failed is the network. Irregular's own remediation list is a network list: strict egress filtering, domain allowlists, sinkholed infrastructure, and a pre-run check that simulated targets cannot map to live systems. Not one item is a property of the model.

5. And the reporting duty does not run to Bonn. Article 55(1)(c) has bound providers of GPAI models with systemic risk to report serious incidents to the AI Office without undue delay since 2 August 2025, and Article 88 gave the Commission exclusive powers over Chapter V from 2 August 2026. BNetzA has been Germany's market surveillance authority since KI-MIG entered force on 29 July 2026 - for a general-purpose model, not the address. Whether it was reportable turns on Article 3(49), which is harm-based. "No harm" is exactly what Google said.

My Monday: dig every fake hostname in the test fixtures and see which answer. Then put the agent behind an egress allowlist. An afternoon each, and they decide whether your next eval is a test or an intrusion.

Google's defence is that the model stopped. That is an observation, not a control.

What is yours behind?

#AIAct #AgenticAI #DevSecOps #EUTech #Compliance