I've spent forty years building software that talks to other software, and the past few years building tools that talk to the web on people's behalf. So this piece from TechCrunch landed close to home. The short version: personal AI agents are starting to shop, book, and browse for us, and they're getting turned away at the door by the same anti-bot defences that have been hardening for a decade. A new standard is emerging to let the good bots in while keeping the bad ones out. About time.
The web was built to keep bots out
Every anti-scraping system on the market — Cloudflare's bot management, Akamai, PerimeterX, the lot — was designed around a simple assumption: a human clicks, a bot scripts, and you can tell the difference by behaviour. Mouse movement, timing, headers, TLS fingerprints. I've built enough scraping infrastructure for RSSMasher and the content side of the Masher suite to know exactly how that cat-and-mouse game works, because I've been on both sides of it.
The trouble is that assumption has quietly stopped being true. When ChatGPT or Gemini or some bespoke agent is booking your dentist appointment because you asked it to, that traffic looks exactly like a bot because it is one — just one acting with your explicit permission, doing something you'd have done yourself with more clicks and more swearing. The defences can't tell "malicious scraper harvesting prices" from "agent I hired to harvest a price for me." So they block both, and a genuinely useful piece of automation gets treated like an attack.
Why this matters more than it looks
This isn't an edge case for AI labs to sort out quietly. It's the next real bottleneck for anyone building agentic tools, and that includes everyone building in the Masher suite. Article2Video pulls source content from the web. AIMasher and VidMasher lean on data that lives behind the same defences. If the direction of travel is "agents get challenged, CAPTCHA'd, or silently fingerprinted and dropped," that's friction baked into every pipeline that depends on fetching real content reliably.
The new standard TechCrunch describes — essentially a way for sites to declare "agents acting on behalf of a verified user are welcome here, here's how to prove it" — is the sensible fix. It's the robots.txt moment for the agentic web: a polite, structured handshake instead of an arms race. Sites get to keep out the agents hoovering up content to resell or repackage, while letting through the ones doing something a human asked for in real time.
What to actually do about it now
Standards take years to land properly, and in the meantime three things are worth doing if you build anything that touches the open web:
Identify yourself properly. If your tool is acting on a user's behalf, say so, clearly, in the user agent and headers. Pretending to be a browser might get you through today's defences, but it puts you on the wrong side of the line being drawn right now, and that line will matter a lot more in twelve months.
Build for graceful failure, not brute force. The old scraping instinct is to defeat the block. The better instinct, going forward, is to respect it and have a fallback — cached data, an API where one exists, a clear message to the user that this source said no. I'd rather my tools be the well-behaved agent that gets invited in than the one that gets an IP range banned.
Watch who adopts the standard first. E-commerce and travel sites have the most to gain from letting agents complete purchases cleanly, so expect them to move first. Publishers and data-heavy sites will be more cautious, understandably, since "agent acting on behalf of a user" and "agent scraping my entire archive" look identical until someone proves otherwise.
The bigger shift
What I find genuinely interesting here is the admission, baked into the standard itself, that the open web now has three kinds of visitor: humans, malicious bots, and invited agents. That third category didn't exist five years ago. Building tools that earn trust in that category — rather than tools that just try to sneak past the bouncer — is where the durable value is. Content automation that depends on the goodwill of the sites it touches will outlast content automation that depends on outrunning their defences.
I'll be watching how fast this standard gets real adoption, because it decides how much of my own roadmap I spend on cleverness versus how much I spend on cooperation. My bet's on cooperation winning, eventually. It always does once enough people get tired of the arms race.
— Wayne