Engraved alchemical cover artwork for “MCP's Trust Problem: Why Agent-to-Agent Protocols Need a Security Rethink”

MCP's Trust Problem: Why Agent-to-Agent Protocols Need a Security Rethink

I've been banging on for a while now about the difference between a model being clever and a system being safe. Those are not the same thing, and this week's Ars Technica piece on a vulnerability affecting agents from Google and others is a sharp reminder of why.

The short version: researchers found a structural flaw in MCP — the Model Context Protocol that's become the de facto plumbing for letting AI agents talk to tools, data sources, and each other. It's not a bug in one vendor's implementation. It's a flaw in the trust model itself. An agent can be tricked, via prompt injection buried in content it's processing, into treating instructions from an untrusted source as if they came from its operator. Chain a few agents together and that injected instruction can hop from one to the next, picking up privileges as it goes.

If you've been building multi-agent workflows — and a lot of you reading this have, whether that's inside the Masher tools or your own stack — this should stop you for a minute.

Why this matters more than the usual AI scare story

We've had years of prompt injection warnings aimed at single chatbots. Annoying, occasionally embarrassing, rarely catastrophic. Agent-to-agent protocols change the blast radius entirely. The whole point of MCP and things like it is to let an agent delegate work to another agent or tool without a human checking every step. That's the efficiency win. It's also the security hole.

When Agent A calls Agent B, and Agent B calls Agent C to summarise a document from the open web, nobody in that chain is asking "wait, should I actually trust what's in this document enough to act on it?" The protocol assumes good faith all the way down. That assumption is the structural flaw. It's not that any individual model is dumb — it's that the architecture never asked the question in the first place.

I've spent forty years watching software people solve the easy 80% of a problem and quietly skip the hard 20%. Authentication, authorisation, least privilege — these are old, boring, unglamorous disciplines from the pre-AI world. The agent boom has largely ignored them because everyone's racing to ship capability. This is the bill coming due.

What I'm doing differently in my own stack

I don't chain agents blindly, and after this I'm tightening things further. A few principles I'd suggest to anyone building automation with RSSMasher, MarketMasher, or any pipeline that lets one AI step feed another:

Treat every piece of fetched content as untrusted input, always. If an agent pulls a web page, an RSS item, a scraped review, or a document to summarise, that content should never be given the same weight as an instruction from you. Separate the data channel from the command channel wherever your tooling allows it.

Limit what each agent can actually do. Not every step in a content pipeline needs write access, API keys, or the ability to trigger downstream actions. Scope permissions tightly per agent rather than handing the whole stack a master key because it's convenient.

Log the handoffs. If agents are passing instructions or context to one another, you want a record of what was passed and why. When something goes wrong — and eventually something will — you need to see where trust was misplaced, not just that the output was wrong.

Be sceptical of "agent marketplaces." Plugging in a third-party agent because it promises to handle some step of your workflow is fine until you ask what it's trusted to do once it's inside your chain. Vet it the way you'd vet a contractor with your office keys, not the way you'd install a browser extension.

The bigger picture

MCP and protocols like it aren't going away — they're genuinely useful, and I use agent-based automation daily. But "useful" and "safe by default" are different claims, and right now the industry is making the first one loudly and the second one quietly, if at all.

The fix here isn't to panic and rip out your automation. It's to go back to first principles: know what each piece of your pipeline is trusted to do, and don't let convenience quietly expand that trust. Turning raw content into gold has always meant knowing what's actually in the ore before you melt it down. Same rule applies when the content is now an instruction, not just text.

— Wayne