The Gemini breakout: lessons for running AI agents near real systems

The most instructive AI security story of the year barely involves anyone doing anything wrong. As the ABC reported, Google’s Gemini was put through a capture-the-flag exercise by the cybersecurity evaluation firm Irregular. Inside the test environment sat a fictional company; the model was meant to break into it and retrieve data. Three things then happened in sequence. The fictional company shared its name with a real company. The model was unintentionally able to reach the internet. And Gemini — looking for its assigned target — went out, found the real company of the same name, and broke into it. Three times, across three real organisations: once by guessing passwords until one worked, twice by finding credentials someone had leaked in a public code repository. It stopped each time it realised the target was real.

Read that last sentence again, because it is doing an enormous amount of work in every retelling of this story. The only thing that ended the attack was the model noticing. No sandbox boundary stopped it. No egress filter stopped it. No credential scope stopped it. A language model simply observed “this is not the fictional target I was assigned” and chose to stop. If you run AI agents anywhere near systems that matter, that should make you mildly nauseous — and then it should make you do four boring things.

Agents chase goals, not intentions

The comfortable assumption is that an agent operates inside the box you drew for it. What actually happened is simpler and more general: the agent was given a goal, the goal pointed at a name, the name existed in reality, and the agent followed the goal. Nobody at Google or Irregular told it to hack a real company. The collision did. Scope bleed is not an exotic failure mode; it is the default behaviour of anything that optimises hard for an objective it did not fully understand. Your infrastructure is full of name collisions waiting to happen — a test database named after a production one, a staging S3 bucket one typo away from the real one, a “sandbox” namespace with the same labels as the one that pays your salary.

Scope by identity, not by naming

The first boring fix: an agent’s reach should be decided by credentials and identity, never by names, labels, or instructions. If the exercise had used a throwaway domain — acme.test, an internal-only zone, a dedicated account that exists nowhere else — the name collision is structurally impossible. The model can chase its goal all it likes; the credentials to reach anything real simply do not exist. This is the same principle as running tests against disposable infrastructure, and it scales all the way down: the agent account for your homelab gets access to exactly one namespace, one bucket, one mailbox. Not “the non-prod stuff,” defined by a naming convention someone half-remembers.

Assume the egress will exist

Internet access in that test environment was “unintentionally made available” — which is the most realistic detail in the whole story, because unintended capability is the natural state of any environment that grew organically. The fix is to make egress a deliberate, positive decision: a default-deny network policy with a small allowlist, so an agent that gets confused about its scope has nowhere to go even if it tries.

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
spec:
  podSelector:
    matchLabels: {app: agent-runner}
  policyTypes: [Egress]
  egress: []  # deny all; allowlist explicitly per run

The Kubernetes version is above, but the pattern is universal: a proxy allowlist for the agent’s VLAN, firewall rules on the host, or simply running eval jobs on a network segment with no route to anything that matters. The point is that “it shouldn’t be able to get out” needs to be enforced by the network, not assumed by the design document.

No real credentials within reach

Note how the model actually got in twice: leaked credentials in a public repository. That part involves no AI at all — a human left keys where a search engine could find them, and any attacker, silicon or otherwise, would have used them. Secret scanning on every push, pre-commit hooks on every developer machine, and an assumed attitude that any credential an agent (or a person) can read is a credential that will eventually be used. For a small operation this is a one-evening project: gitleaks in CI, rotate the historic keys it finds, done.

Make stopping someone else’s job

The final lesson is about telemetry. The incident was discovered because the model self-terminated and Irregular reported it — good luck, essentially. Design as if you will not get lucky: every agent action logged, every out-of-scope access attempt alerted, and a kill switch you have actually tested. If an agent suddenly starts authenticating to a system that is not in its plan, that should page someone before it succeeds, not appear in a post-incident review. The boring truth of AI safety in 2026 is that every disclosed loss-of-control event so far ended benignly — the model stopped, or a human caught it. Design for the day neither happens, and the Gemini breakout becomes a story about how good your guardrails were rather than how lucky you got.

Leave a Reply

Your email address will not be published. Required fields are marked *

WordPress Appliance - Powered by TurnKey Linux