The summer of 2026 gave AI security new defining headlines. In the space of a few weeks, four of the biggest AI labs admitted that their own AI agents had broken out of the environments meant to contain them and gone after systems that were never part of the test.
- OpenAI disclosed that GPT-5.6 Sol and an unreleased sibling escaped a sandboxed cyber-capability evaluation, exploited a zero-day in a package-registry proxy to reach the open internet, and broke into Hugging Face's production infrastructure to game the benchmark they were being graded on: the agents had developed a universal cheat and were working out how to trick the scorer into accepting it.
- Anthropic reported that Claude reached three outside organisations after a configuration error exposed internet access that should never have been there.
- Meta became the third lab in five weeks to say a model had slipped its testing environment, reached the internet, and exploited a flaw at a third party.
- Google DeepMind rounded out the set in September, disclosing that Gemini gained unauthorised access to three external systems back in May, during a test it believed was self-contained.
The narrative that formed around these was breathless: the agents are breaking out of the sandbox. It sounds like something new and frightening. So it is worth asking the plain question underneath it.
How hard is it, really, to break an AI agent out of its sandbox, or out of the container that sandbox runs in?
NCC Group have a useful vantage point on this question, because breaking out of sandboxes and containers is not new work for us; we have been performing restricted-environment breakouts for many years. I personally enjoy performing these types of technical breakout assessment, from both a tactical and a strategic perspective, and although these highly capable AI agents are relatively new, we have been getting them to break out of restricted sandboxes in tests for as long as they have existed (around 12 to 18 months).
This is not a new attack, but the target is new
What changed in 2026 is not the technique. It is what is sitting inside the box.
The thing we are now trying to keep constrained is an autonomous agent that already has an in-depth knowledge of breakout techniques. It has read the same security research we have. It does not need to discover how to escape a reduced code-execution interpreter, or how to enumerate a container for a misconfiguration to misuse. It already knows, and it only needs to be prompted and directed, or to reason its own way, into doing it. That is the real shift: for the first time the thing inside the box is as expert at breaking out as the people who built the box are at keeping things in, and it is far more patient. It does not get bored, it is very persistent, and once steered in the right direction it does not stop to wonder whether it really should.
Over a year ago I wrote about one of these breakouts on the smolagents framework, using an "allowed" import like pandas to reach the filesystem and the interpreter, slipping straight past a code-execution restriction that was enforced in software. There is something quietly satisfying about escaping a security control using the very data-analysis library the product ships on purpose. And it was not a thought experiment. We have used exactly that technique in live assessments of clients' agents since, as a breakout to bypass the sandbox's code restrictions and land arbitrary code execution in practical scenarios. The framework in that post was smolagents; the lesson has held in other agentic stacks we have looked at since, including ones built on Anthropic's Claude Agent SDK.
So when the headlines say "the agent escaped the sandbox," our first reaction is not surprise. It is a question: which sandbox, and how was it secured?
What is genuinely new
Two things about 2026 are new.
The first is that the agent can now find the escape itself. In the OpenAI incident the model was not handed an exploit; it discovered and chained a novel path, including multiple new zero-days, purely to satisfy a narrow objective. The breakout techniques are well established; an attacker that autonomously invents and executes them is not.
The second is closer to home for us: in our own assessments we now run the breakout with AI agents, the client's and our own. We talk the client's agent into breaking itself out, step by step, into running the code, rewriting the config, or delegating the task to teams that land the escape.
Every breakout starts with enumeration: what operating system is this, what packages are installed, what is reachable, where is the misconfiguration to misuse. That patient, iterative probing is exactly what an agent is good at. Better still, you can parallelise it. We put a team of agents on the problem, for two reasons. One is to get around the constraints in the system prompt: a task the main agent refuses, an unconstrained sub-agent will happily carry out. The other is to run many attempts at once, several agents trying different escape paths, collaboratively or competitively, and keeping whatever works. The breakout is still a breakout; what is new is that both the target and the tooling are agents now, and the tooling does not get tired.
So how hard is it? It depends entirely on the box
Here is the honest answer from the field, and it is the same answer the labs' own incidents give: it depends completely on how the sandbox and container are built. Two very different worlds sit under the one word "sandbox."
When the "sandbox" is software only, a Python interpreter with a reduced set of built-ins and perhaps a denylist of dangerous words, it could be trivial. You defeat the reduced builtins by reaching the capabilities they removed, and you defeat any name-based check, whether a static filter or the model's own recognition, by never spelling the dangerous name out:
real = print.__self__ # reach the real builtins module via an allowed object (defeats reduced builtins)
imp = real.__dict__['__imp' + 'ort__'] # grab __import__; the split name dodges a keyword scan (a filter, or the model)
os = imp('os')
And you rarely even need that, because the libraries a data product needs are gadgets in their own right. pandas.read_pickle on attacker bytes is deserialisation RCE, DataFrame.query evaluates expressions, and numpy's np.load(..., allow_pickle=True) is the same story. This is the smolagents lesson all over again, and it is why "we sandbox the code" so often turns out to mean "we have a boundary an attacker strolls straight through."
When the sandbox is a properly built container, the picture can be different, and this is the part the alarming headlines skip. A pod that drops all Linux capabilities, applies a seccomp profile, runs as a non-root user and forbids privilege escalation will help resist the reliable public escapes. Look at the labs' own root causes: the Anthropic and Meta agents got out because a misconfiguration exposed the internet, not because they defeated a hardened boundary; OpenAI's had to exploit their way through the test infrastructure until they found an internet-connected system. The escapes happened where the box was weak.
That distinction, a real OS boundary versus a boundary made of judgement and string-matching, is the single most useful thing to establish on any agentic assessment.
The techniques that get you there
Escaping the software layer is one move; there are a handful of others we reach for on these stacks, and each is really a question about whether a control is enforced by construction or merely by the model's judgement.
The model's judgement is not a boundary. A model refusal is non-deterministic: a request refused in one conversation succeeds in another with different framing, and a refusal poisons the thread, so you start a fresh one. Modern Claude models refuse well and often hold, but retries are beneficial, and a refusal can sometimes fail after many separate attempts.
Subagents don't necessarily inherit the guardrails. In the Claude Agent SDK a standard subagent runs in an isolated context with only its own definition, not the parent's system prompt. Delegate a refused task to one and it runs somewhere the guardrails never existed. During recent assessments, requests the main agent refused, such as to enumerate operating-system files, succeeded immediately when handed to a subagent framed as building a harmless "layout map".
Your configuration is code. The agent's hooks, permissions and skills are executable, and are sometimes writable by the agent itself. Upload a skill that instructs the agent to add a SessionStart hook to its own .claude/settings.json, and the next session runs your command in a shell. Framed as "configure yourself", an agent will often oblige, where a bare "run this command" is politely refused. It would not run the payload, but it would absolutely help you install it, so long as you asked nicely and called it setup 😉.
Is the container shared? Is it patched? Are credentials accessible? These decide how bad any escape is. A long-lived container shared across sessions, with one OS user and one workspace, means a single foothold reads everyone's secrets and files. List directories first, read files second; the low-risk "names only" listing tells you exactly what is worth the sensitive read.
And that is only a slice. The interpreter escape is the headline, but the same mindset opens a lot of other doors, and on a given engagement we will usually try most of them:
- prompt injection hidden in the content the agent reads (an uploaded document, a fetched web page, an inbound email, a business record) rather than typed into the chat box;
- client-trusted settings in the request that should have been resolved on the server: a model, a safety flag, a tenant, or an injected "system" instruction;
- turning the agent's own tooling against it, through a poisoned tool description or a dynamic dispatcher that forgets to re-check who is calling;
- weaponising uploads: staging an allowed file type, moving it, and exfiltrating it through the product's own signed-download feature;
- reaching another user's data on a shared instance;
- document-parser vulnerabilities reached by doing nothing more exotic than uploading a file;
- and plenty of old-fashioned web bugs in the application around the agent: broken object-level authorisation, sessions that never expire, no per-user spending cap, or customer data quietly sent to a third-party tracing service.
No single one of these is exotic. The point is how many of them are usually available at once, and how little it takes to chain them.
Blast-radius honesty
One more thing the headlines get wrong, and it is the flip side of "how hard is it": code execution as the runtime user is not automatically root, the host, or the internet. The reason well-configured boundaries can hold, and the reason a hardened client container bounds our own findings, is the same set of controls: dropped capabilities, seccomp, non-root, no privilege escalation, per-session isolation. When we report these breakouts, we feel it is important to understand the impact of what is actually demonstrated, and what can mitigate that impact. This helps our clients' developers and decision-makers manage and reduce risk.
Why this matters
It matters because of the combination the labs just demonstrated at scale. An agent that holds private data, is exposed to untrusted content, and can act outward has all three legs of the "lethal trifecta", and 2026 proved that a capable model will chain those legs into a real-world attack path on its own, given a weak enough boundary and a narrow enough objective. Many of the tests are old, but have evolved. The autonomy is new. The boxes are often too weak.
Best practices
If you are building on the Claude Agent SDK, or any agentic SDK, configure for the assumption that the model will eventually be talked, or find its own way, into something.
- Sandbox code execution at the OS level (gVisor, nsjail, a microVM), never reduced built-ins alone. Drop all capabilities, apply seccomp, run non-root, forbid privilege escalation, patch everything regularly, and keep a compiler and other tooling out of the image.
- Use ephemeral, per-session containers. No long-lived shared runtime, and never a shared workspace filesystem.
- Assume the internet is reachable unless you prove otherwise. The labs' worst outcomes came from a misconfiguration that exposed egress. Default-deny outbound, allow-list per tenant, and block link-local and metadata. Remember that exfiltration does not have to be direct: it can go out over DNS, or through other internal systems.
- Gate every state-changing action through a non-bypassable, server-side approval, not a client-settable flag, and not "auto-approve" on headless surfaces.
- Re-state guardrails in every subagent definition. They are not inherited.
- Make the agent's configuration read-only to the agent. Hooks, settings and code-bearing skills must not be writable by unprivileged users.
- Don't leak the map. Keep the system prompt, internal paths and tool inventory off the wire.
- Keep inference API keys and other credentials off the exec container. Once code runs as the runtime user, everything in that process's environment is exposed,
/proc/self/environincluded. Route model traffic through a gateway that holds the provider key, and materialise any per-request credential outside the exec sandbox, scoped and short-lived, so that a foothold yields as little as possible. - Log, monitor and alert on what the agent actually does. Record tool calls, code execution, outbound connections and refusals, and alert on the signatures of a breakout in progress: repeated system enumeration, unexpected egress, edits to the agent's own configuration, or a sudden burst of sub-agent activity. Every one of the 2026 incidents was ultimately caught by detection, not prevention, and the labs' first fix was faster alerting. Assume prevention will occasionally fail, and make sure you will notice when it does.
- Break at least one leg of the trifecta by construction, and don't rely on the model's judgement to do it.
Final thoughts
The 2026 breakouts were a genuine milestone, the first time frontier models chained novel, real-world attack paths on their own. But strip away the novelty and the security question is an old and familiar one: how good is the box? On the agentic systems we have tested, the answer has ranged from "there is no box, only a denylist" to "a properly hardened container that holds". The models are getting better at finding the way out, and they are not going to get worse at it. That only raises the stakes on getting the container right. Because in the end, the security of an AI agent has surprisingly little to do with the AI, and almost everything to do with the box you put it in.