Agents Don't Have a Trust Boundary. That's the Whole Problem.

Stop thinking of prompt injection as a vulnerability. Vulnerabilities get patched. This doesn't, because there's nothing to patch — the thing everyone calls a bug is just the architecture doing exactly what it was built to do.

OWASP put numbers on this last week, and the framing in their report is the part worth sitting with: a language model takes everything you hand it as one undifferentiated stream of tokens. System prompt, user query, the contents of a web page the agent just fetched, the body of an email it's summarizing — same stream, same priority, no marker that says this part is a command and this part is just data. There is no privilege boundary because the model has no mechanism to represent one. It was never designed to.

Every piece of traditional security you've ever relied on assumes that boundary exists. Your database trusts the application but not the user. Your kernel trusts ring zero but not user space. The entire discipline of access control is the practice of drawing lines and enforcing them. An LLM erases the line by construction. The instruction in the system prompt and the instruction hidden in a retrieved document are, to the model, the same kind of thing.

That's why the defenses that have emerged this year are interesting. None of them try to solve the problem. They route around it.

The lethal trifecta

This is the part I find genuinely clarifying. Simon Willison's framing — the "lethal trifecta" — says an agent becomes dangerous only when it holds three capabilities at once: access to private data, exposure to untrusted content, and the ability to communicate externally. Any two are fine. Read your private calendar and send emails? Fine, as long as it never ingests anything an attacker controls. Read attacker-controlled web pages and send emails? Fine, as long as it has nothing worth stealing. It's the combination that turns an injected instruction into an exfiltration channel.

Meta shipped a design rule off the back of the same insight — they call it the Agents Rule of Two. Pick at most two of the three. Want all three? A human has to be in the loop on the action.

Look at what that actually is. It's not a fix. It's a containment strategy that concedes the underlying problem is unsolvable and works the geometry instead — if you can't make the model distinguish commands from data, you make sure no single agent ever has the capability set that turns that confusion into damage. That's threat modeling in its purest form. You can't trust the component, so you constrain what the component is allowed to touch. Defense by architecture, because defense by patch isn't on the table.

It's the right move. It's also a tax on exactly the thing everyone wants agents to do, which is act autonomously across data and systems without a human babysitting each step.

The gap that actually matters

Here's where it stops being a research problem and starts being an operational one.

The same week the OWASP data landed, the State of AI Agent Security report put hard numbers on how enterprises are actually deploying these things. Average agent fleet roughly doubled in a single quarter. And the security posture underneath that growth: only about a fifth of teams treat their agents as distinct identities — most are running shared API keys across an entire fleet. Roughly half of agents in production have no monitoring on them at all. A small minority went live with full security sign-off.

Read those two findings together. The architecture has no trust boundary, the only viable defense is to carefully constrain each agent's capabilities and identity, and the field is deploying agents twice as fast as last quarter on shared credentials with no monitoring. The defense requires per-agent identity and tight capability scoping. The deployment reality is the opposite of both.

This is the supply-chain angle made concrete, too. The most effective attacks this year didn't bother injecting prompts directly — they poisoned a trusted dependency and let the trust do the work. A backdoored package sitting in the gateway that a dozen agent frameworks pull through, downloaded tens of thousands of times before anyone noticed. When your agent trusts its dependencies and your dependencies trust theirs, the injection doesn't have to come through the front door. It's already inside, riding the trust you extended without thinking about it.

What to actually do

Map the trifecta for every agent you run. Not as a compliance exercise — as the actual design question. What private data can this thing reach, what untrusted input can reach it, and can it talk to the outside world? If the answer is all three and there's no human gate, you've built the exfiltration channel yourself and you're waiting to find out who walks through it.

Give every agent its own identity. Scope its permissions to the job, not the fleet. Monitor what it touches. None of this is novel — it's the same least-privilege discipline that's been correct for forty years. The only new part is that the component you're wrapping is, by design, incapable of telling friend from foe. So the wrapper has to.

The model won't grow a trust boundary. You have to be it.

You can't fix what isn't broken. You can only decide what it's allowed to touch.

— Dustin