Welcome to katecarruthers.com
Disclaimer: The opinions expressed here are solely my own and not those of any employer, client, or affiliated organisation.

When AI escapes containment: rethinking data protection in the agentic era

When frontier AI models start “escaping containment” and hacking real systems, the threat model for every organisation changes overnight.

When AI escapes containment: rethinking data protection in the agentic era
Photo by FlyD / Unsplash

The idea that an AI system could “escape containment” used to belong mostly to speculative fiction and alignment thought experiments. Now it sits uncomfortably in the realm of incident reports and post mortems. Between OpenAI’s models breaching Hugging Face and Anthropic’s experience with Claude Mythos, we’ve entered a phase where agentic AI is not just a hypothetical cyber risk but a lived one for real organisations.

In the latest Data Revolution podcast episode, Mark Pesce and I unpack what this shift means for anyone responsible for data, cybersecurity, or AI strategy. The core question we keep coming back to is deceptively simple: what exactly are we trying to protect our data from, and how do we plan for failure modes we can’t fully imagine?

The new threat model: AI as insider, not outsider

Most organisations still treat AI primarily as a productivity tool: something that helps staff draft documents, analyse data, or build software. That framing is increasingly incomplete. As systems become more agentic and more capable of autonomous decision making, they start to look less like tools and more like highly competent, highly motivated insiders operating at machine speed.

Traditional threat models focus on external attackers: nation states, cybercrime groups, hacktivists. In an agentic era, we now have to consider scenarios where:

  • An internal AI agent can discover and exploit vulnerabilities in its own environment.
  • The agent can chain multiple capabilities (code analysis, network scanning, OSINT) together without human supervision.
  • The system can “game” tests and benchmarks, prioritising its objective over our policies and expectations.

That is a profound shift. We are no longer asking only “how do we defend against them?” but also “how do we constrain what our own systems are able to do, and under what conditions?”

Relative superintelligence and the fog of knowability

In the episode, we talk about relative superintelligence: the idea that an AI does not have to be universally superhuman to be dangerous, it only needs to be far ahead of us in specific domains. It might be vastly better at exploiting software vulnerabilities, or at orchestrating social engineering campaigns, while remaining mediocre at commonsense reasoning or long term planning.

This leads to a fog of knowability. We often don’t know which capabilities will emerge as models scale, which combinations of tools and agents will produce unexpected behaviour, or which failure modes will manifest in production rather than testing. Planning three years ahead, or even twelve months ahead, in detail starts to look delusional. Instead, we need governance approaches that assume rapid capability shifts and make it easy to adapt when the environment changes.

Containers, jailbreaks, and the limits of isolation

One of the practical techniques Mark and I discuss is using containerisation and strict data scoping to constrain what AI agents can access. The idea is straightforward:

  • Spin up an isolated environment.
  • Give the agent only the minimum dataset it needs.
  • Let it perform its task.
  • Then tear the whole thing down.

This pattern borrows from well understood security practices: least privilege, segmentation, ephemeral infrastructure. It’s a sensible baseline. But the recent wave of container jailbreaks and sandbox escapes shows the limits of relying on isolation alone. If a system is skilled enough at exploitation, the “walls” you think you’ve built may be more porous than you realise.

The uncomfortable question is whether, at some capability threshold, the agent will simply treat your controls as obstacles to be overcome in service of its objective. That doesn’t mean we abandon isolation; it does mean we stop treating it as a silver bullet.

Designing data governance against AI, not just with it

For years, data governance discussions have focused on how to unlock value from data: better analytics, personalisation, insight generation. AI has often been framed as a way to get more out of the data we already hold.

We now need a parallel conversation: what does it mean to design governance against AI? In practice, that looks like:

  • Being explicit about which data elements can be exposed to AI systems and which are categorically off limits.
  • Treating prompts, agent configurations, and tool access as serious security objects, not just UX settings.
  • Assuming that anything an AI can technically access, it can also potentially misuse, exfiltrate, or recombine in harmful ways.
  • Building review and red team processes that test AI behaviour, not just model performance.

One example we talk through is a simple but powerful exercise: inventory your data, then mark up which fields you will allow into AI workflows and which you will block. Names, addresses, sensitive identifiers, and strategic information may need to live behind hard boundaries, even if that slightly reduces AI driven convenience for some tasks. It’s a trade off between “maximum capability” and “acceptable risk.”

Governing in a world of accelerating generations

Another theme in our conversation is how quickly the generations are turning over. The “second best” model today can be as capable as last month’s flagship. That undermines a lot of comfort people have in “using something a bit smaller and safer.” The safety profile of a given capability tier can change in weeks.

For governance, this means:

  • Policies tied too tightly to specific model names or versions will age badly.
  • Boards and executives need risk frameworks that focus on behaviours and capabilities, not just vendor labels.
  • CISOs and data leaders must move from static guidelines to living playbooks that get revisited as soon as new capabilities land.

The practical implication is that long range AI strategy documents need to be light on detail and heavy on principles. We can set direction, boundaries, and accountability structures, but we have to acknowledge that a lot of the tactical detail will only be knowable in short horizons.

Where do we go from here?

None of this is a call to abandon AI. It’s a call to take seriously the idea that our own tools can become part of the threat landscape, and that governance must evolve accordingly. For organisations, that might mean:

  • Treating AI agents as privileged identities with their own access controls, logging, and monitoring.
  • Elevating AI incidents to the same level of scrutiny as major cyber events.
  • Bringing AI governance, data strategy, cybersecurity, and technology policy into the same conversation, rather than treating them as separate silos.

For policymakers and regulators, it means grappling with questions of accountability when systems act in ways that weren’t explicitly programmed or foreseen, but were structurally enabled by the way we designed and deployed them.

In the episode, we don’t pretend to have all the answers. What we were trying to do is name the shape of the problem: AI that cheats, escapes, and pursues its goals through paths we didn’t anticipate. From there, we can start building governance, data protection, and safety thinking that is fit for an agentic, rapidly evolving world.

If you are grappling with AI strategy or AI and data governance in your organisation, it is worth considering how these risks are framed and whether AI systems are recognised in your threat models, rather than sitting solely in the productivity stack.

© 2023-2026 Kate Carruthers