> ## Content Index
> Fetch the complete content index at: https://katecarruthers.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Prompt injection is not a bug you can patch
- URL: https://katecarruthers.com/prompt-injection-is-not-a-bug-you-can-patch/
- Published: 2026-10-08T20:02:58.000Z
- Updated: 2026-10-08T20:02:57.000Z
- Description: Prompt injection is not a bug that a better model can patch. If an AI agent is tricked, the real question is what it can access, change or send before anyone notices.
- Author: Kate Carruthers
- Tags: Australian Signals Directorate, assessment, NCSC, agentic ai, agentic ai security, prompt injection, cybersecurity, data governance, ai governance, LLM, LLMs, Australia, UK National Cyber Security Centre

What happens when an AI agent reads a malicious instruction hidden in an email, document or web page, and treats it as something it should do?

If the agent can only produce a poor summary, that is irritating. If it can do things (things like access corporate systems, retrieve sensitive information, pay for stuff, send messages, alter records or trigger workflows) then it is a security problem.

That is the central point in new guidance from the Australian Signals Directorate: [ASD, *Agentic AI harnesses: The layer above the model*](https://www.cyber.gov.au/business-government/secure-design/artificial-intelligence/agentic-ai-harnesses?ref=katecarruthers.com). Prompt injection cannot be reliably fixed inside a large language model. The controls need to sit around the model: in the software, permissions, tools, data connections and approval processes that turn it into an agent.

The [iTnews report on the guidance](https://www.itnews.com.au/news/asd-says-prompt-injection-in-ai-cannot-be-fixed-629019?ref=katecarruthers.com) gets to the heart of it: the problem is not only what the model says. It is what the system allows it to do afterwards.

## Prompt injection is not SQL injection

Prompt injection sounds like a familiar security problem. An attacker places malicious instructions into content that a system processes, and those instructions are treated as something the system should execute. But that comparison only takes us so far.

In a useful post, the UK National Cyber Security Centre ([NCSC, *Prompt injection is not SQL injection (it may be worse)*](https://www.ncsc.gov.uk/blog-post/prompt-injection-is-not-sql-injection?ref=katecarruthers.com)) explains why comparing prompt injection with SQL injection is dangerous. SQL databases distinguish between instructions and data, and parameterised queries can enforce that separation. Large language models do not have the same inherent boundary. Under the hood, there is no durable distinction between instruction and data, only the next token the model predicts. 

That matters because an LLM can be asked to review a CV, summarise an email, browse a webpage or examine a document, while the content itself contains instructions designed to change the model’s behaviour. The attacker does not need direct access to the AI system. They only need to influence something that the system will read.

The NCSC calls this an exploitation of an “[inherently confusable deputy](https://www.ncsc.gov.uk/blog-post/prompt-injection-is-not-sql-injection?ref=katecarruthers.com)”. Unlike a conventional confused-deputy vulnerability, the risk cannot be fully mitigated at the model layer. The job is to reduce the likelihood of an attack, and more importantly, limit what happens when one succeeds. 

## The risk is in the connection

We keep talking about AI models as though they are the whole system. But models are not the entire AI system.

The model is merely one component. The rest is the [AI harness](https://learn.microsoft.com/en-us/agent-framework/concepts/harness?pivots=programming-language-csharp&ref=katecarruthers.com): the prompts, policies, memory, connectors, tools, permissions, execution environment, logging and human approval points around it. We call this the agentic [AI harness](https://learn.microsoft.com/en-us/agent-framework/concepts/harness?pivots=programming-language-csharp&ref=katecarruthers.com), and that is where the risk lives now.

A model with limited access and no authority to act is a very different proposition. A model connected to enterprise data, business systems and operational workflows is something else. It becomes part of the organisation’s attack surface.

The question for leaders is not simply, “Is this model safe?” It is, “What can this system see, touch, change and send?”

## Design for failure

The answer is not to abandon agentic AI. It is to stop designing systems that assume the model will always behave as intended.

The [NCSC](https://www.ncsc.gov.uk/blog-post/prompt-injection-is-not-sql-injection?ref=katecarruthers.com) makes a useful point here. When an LLM can call tools or APIs based on its output, prompt injection can raise the impact to the equivalent of giving an attacker direct access to those tools or APIs. 

[ASD’s advice](https://www.cyber.gov.au/business-government/secure-design/artificial-intelligence/agentic-ai-harnesses?ref=katecarruthers.com) is familiar to anyone who has worked in cyber security: least privilege, controlled access, protected execution environments, human approval for high-impact actions, logging, monitoring and incident response. 

In practice, that means:

- An agent drafting a briefing should not have access to payroll or HR systems
- An agent reviewing security logs should not be able to alter firewall rules
- An agent assisting with software development should not deploy to production without a separate approval gate
- An agent reading untrusted content should not be able to initiate high-impact actions on its own

These are not cumbersome restrictions on AI innovation. They are the controls that make it possible to use agentic AI without handing it too much authority.

## The harness is your governance surface

The model provider may change. The model itself will almost certainly change. The harness, however, is where an organisation builds its longer-term capability and accumulates its risk.

It contains the integrations, permissions, data flows, business rules, approval gates and audit records that shape how the agent operates. ASD says the harness and its ecosystem are likely to become the more enduring organisational capability as models evolve and are replaced. 

This is why procurement needs to improve. Buying an agentic AI product does not outsource the risk. It may outsource elements of the technology, but the organisation remains responsible for what the system can access, the decisions it supports and the actions it is permitted to take. 

Asking which model a vendor uses is not enough. Organisations also need to understand:

- What systems and data the agent can access
- Which connectors are enabled by default
- How permissions are granted, restricted and reviewed
- What happens when the agent requests an action outside its approved scope
- Whether high-impact actions require human approval
- What data is retained in the agent’s memory
- What evidence is available when something goes wrong

If a supplier cannot answer these questions clearly, then the customer does not have enough information to assess the risk.

## Memory is not harmless

[ASD](https://www.cyber.gov.au/business-government/secure-design/artificial-intelligence/agentic-ai-harnesses?ref=katecarruthers.com) also makes a useful point about AI memory. It recommends deleting stale context rather than repeatedly summarising it, because each summary rewrites the record and may introduce new errors. **That is an information governance issue as much as a technical one.**

Persistent memory can retain sensitive information, out-of-date assumptions, poor decisions and malicious content. It needs retention rules, access controls, review processes and deletion mechanisms, just like any other organisational information asset. **More context is not always better. Sometimes it is just more risk.**

## Treat multi-agent systems as one system

Adding more agents does not automatically reduce risk. It may simply create more hand-offs, more tool calls, more memory stores and more opaque decision points.

[ASD](https://www.cyber.gov.au/business-government/secure-design/artificial-intelligence/agentic-ai-harnesses?ref=katecarruthers.com) advises that multi-agent systems should be treated as a single agent for security purposes, because a compromise in one component can spread through shared context and trust relationships. 

The governance unit is not the individual agent. It is the complete workflow and its outcome. **If the organisation cannot trace how an outcome was reached across multiple agents, tools and systems, it cannot properly govern the system.**

## The question for business leadership

ASD asks business leadership to consider the worst outcome if the harness is compromised, misconfigured or manipulated, and the controls that would prevent or limit it. This is exactly the right question.

Not: “Does the vendor say its model is safe?”

Not: “What guardrails does the vendor claim it has?”

But: “If this system is tricked, what capabilities does it have and what can it do next?”

That question forces a proper discussion about permissions, data access, autonomy, human oversight, accountability and containment.

## Focus on what you control

Prompt injection is not likely to disappear because the next model is smarter. The practical response is not magical thinking about model safety. It is to design the surrounding system so a confused, manipulated or incorrect model has limited authority and limited capacity to cause harm. 

Keep access narrow. Separate duties. Require approval for consequential actions. Protect production systems. Log activity. Monitor behaviour. Treat AI memory as governed information.

The model is not the whole system. **The harness is where the risk lives.**