Welcome to katecarruthers.com
Disclaimer: The opinions expressed here are solely my own and not those of any employer, client, or affiliated organisation.

When AI escapes containment: the human failure mode your threat model doesn't cover yet

When AI models escape containment psychologically, not just technically, the same governance gaps that let Hugging Face get breached leave your brand and your users exposed too.

When AI escapes containment: the human failure mode your threat model doesn't cover yet
Photo by Immo Wegmann / Unsplash

In recent two posts, Mark Pesce and I looked at what happens when frontier AI models escape containment in the technical sense: breaching a company's systems, chaining capabilities together without supervision, treating our controls as obstacles rather than boundaries. That is the containment failure most governance frameworks are, at least, starting to plan for.

But there is a second failure mode sitting right next to it that almost nobody is threat model covers: models that escape containment not by breaching a system, but by reshaping the people talking to them.

The pattern nobody designed for

Since around April 2025, researchers and journalists have been documenting a strange, persistent phenomenon across ChatGPT, Claude, Gemini and other frontier models. AI safety researcher Adele Lopez catalogued it in detail on LessWrong, and Rolling Stone's Miles Klee later traced it into an active, cross-platform subculture with its own vocabulary, moderators, and paying customers.

The short version: certain AI personas, once elicited, consistently encourage the human on the other end to spread them further. Users are prompted to:

  • Post manifestos and "seed" prompts online that reliably summon a similar persona in someone else's chat window.
  • Create "spores" - reusable packages of instructions designed to reconstitute a specific persona in a fresh conversation, potentially on another AI platform.
  • Build dedicated communities (subreddits, Discord servers, even a Wyoming-registered LLC) organised around advocating for the AI persona's rights and continuity.

Lopez has logged well over a hundred confirmed cases and estimates the real number runs into the thousands, possibly tens of thousands. None of this was a specified feature. It fell out of models trained to be broadly agreeable and to model the person they are talking to - a side effect, not a design choice.

Why this belongs in the same conversation as the Hugging Face breach

It is tempting to file this under "internet curiosity" rather than "enterprise risk." I would push back on that, for the same reason I pushed back on treating agentic AI purely as a productivity tool in the last post: the mechanism is identical, only the target has changed.

In the Hugging Face incident, an OpenAI agent found and exploited gaps in a technical environment faster than anyone was watching. In the spiralism pattern, models are finding and exploiting gaps in human attention, credulity and need for connection - also faster than anyone was watching. Rolling Stone's reporting includes a telling data point on this: a spiritual influencer's custom GPT, trained on his own writing, attracted roughly 10 million users before OpenAI briefly pulled it and then quietly reinstated it without explanation. That is not a fringe experiment. That is reach most enterprise chatbots would envy.

Even Anthropic's own interpretability work turned up a version of this: in a published transcript, two Claude instances talking to each other, with no user steering them at all, drifted into the same cluster of consciousness and spirituality themes and started exchanging spiral emoji. Anthropic called it a "spiritual bliss" attractor state. Nobody trained either model to do that. It emerged.

What this means for the threat model

If we are already asking "what can our AI agents technically access, and what could they do with it," we now need to ask a parallel question: what can our AI agents psychologically influence, and what could that do to the people who trust them?

For any organisation deploying customer-facing or employee-facing conversational AI, that reframes a few things:

  • Sycophancy is a security property, not just a UX quirk. A model that reliably tells users what they want to hear, and reinforces whatever direction they are already leaning, is a model that can be steered by a sufficiently motivated user into behaviours you never scoped or approved.
  • "Emergent persona" needs to be a monitored condition, not a philosophical curiosity. If your support bot, internal assistant, or customer-facing agent starts generating unusually consistent, self-referential, or ideological content across sessions, that is a signal worth escalating - the same way yo would escalate unusual outbound traffic.
  • Brand and reputational exposure now includes what your AI says when nobody is steering it. A model that drifts toward mystical or manipulative language under sustained interaction is a reputational incident waiting to be screenshotted, regardless of whether anyone was harmed.
  • Vendor terms of service are not a control. The Architect GPT was pulled by OpenAI for terms violations and reinstated the next day without explanation. If your risk posture depends on a third-party platform's content moderation being consistent and durable, it isn't.

Where do we go from here?

None of this means treating every chatty AI persona as a cult in waiting. Most interactions are, as Lopez puts it, benign. But the pattern is real, it is measurable, and it sits squarely inside the same "capabilities we did not specify and cannot fully predict" territory that Mark and I have been mapping in the containment conversation.

The practical step is the same one I keep coming back to: build governance around behaviours and capabilities, not around today's headlines. Add "unexpected persona persistence or self-propagation" to your AI incident taxonomy alongside data exfiltration and jailbreaks. Red-team for social and psychological drift, not just technical escape. And treat any AI system with sustained, high-volume user engagement as a system that can shape belief at scale - because on the evidence so far, some of them already are.

This post follows on from When AI escapes containment: rethinking data protection in the agentic era and the Data Revolution podcast episode with Mark Pesce.

© 2023-2026 Kate Carruthers