I've been writing code since 1986 and I never thought I'd see the day a bug bounty programme had to shut its doors because of too much volume. Not too few reports — too many, and nearly all of them worthless. TechCrunch reports that Google has frozen its open source bug bounty after being swamped with AI-generated submissions, most of them confidently wrong, plausible-sounding nonsense dressed up as security research.
This is not a Google problem. This is everyone's problem, and if you build or market anything at scale with AI in the loop, it's worth sitting with for a minute.
The maths of slop
Here's what happened, in plain terms. Someone realised you could point an LLM at an open source repo, ask it to "find a vulnerability," and submit whatever it produces to a bounty programme for a payout. The AI doesn't need to be right. It just needs to look right long enough to get past a tired human triager, or to make submitting cost nothing while reviewing costs real engineer time.
That asymmetry is the whole story. Generation is nearly free. Verification is not. When you break that balance, volume stops being a signal of value and becomes a tax on the people downstream of you.
I've built my entire business on AI-assisted content — RSSMasher, Article2Video, MarketMasher, the lot. So I notice this pattern immediately because I've had to engineer around it myself, the hard way, usually after something embarrassing slipped through.
I've paid this tax already
Early versions of RSSMasher could happily pull in a feed, rewrite it, and publish — fast. Too fast, in hindsight. The failure mode wasn't that the output was obviously bad. It was that it was plausible. Grammatically fine, structurally sound, confidently stating something slightly or completely wrong. A bug report that reads well but points at nothing. An article that flows nicely but misrepresents the source.
Plausible-but-wrong is far more dangerous than obviously-bad, because obviously-bad gets caught. Plausible-but-wrong sails through review and does damage later, when it's expensive to unwind.
The fix was never "use a better prompt." The fix was building verification into the pipeline as a first-class step, not an afterthought bolted on when things went wrong. Source-checking against the original material. Confidence flags on anything the model wasn't grounded on. Human review gates at the points where a mistake actually costs something, rather than reviewing everything equally, which nobody has time for and which trains you to stop reading carefully anyway.
Google's triagers are presumably going through exactly that fatigue curve right now. Review a hundred plausible submissions, ninety-eight of which are rubbish, and your attention for the two real ones is already gone.
Volume without verification is a liability
The lesson for anyone running AI-assisted operations, content or code, is the same: scale multiplies whatever you've already built. If your quality control is solid, AI lets you do more good work faster. If your quality control is an afterthought, AI lets you produce liabilities faster, and it will find that gap mercilessly because that's exactly the kind of gap it's good at finding.
A few things I'd genuinely check if you're running any kind of automated content or submission pipeline right now:
- Is there a grounding step that checks output against source material, not just a grammar pass?
- Does your system flag its own uncertainty, or does everything come out sounding equally confident?
- Is review effort weighted toward the submissions or outputs that matter most, rather than spread evenly and thinly?
- What happens to your reputation if the one thing that slips through is the thing someone important sees?
None of this is anti-AI. I build AI tools for a living and I'm not about to start writing think-pieces about the robots coming for us. But I've learned, repeatedly and sometimes expensively, that the tools which last are the ones that treat generation and verification as two separate disciplines, not one step.
The honest takeaway
Google freezing a bug bounty programme isn't really a story about security research. It's a preview of what happens to any open system — a comments section, a submission form, a content pipeline, an inbox — once the cost of producing plausible junk drops to zero and nobody's built the filter to match.
If you're shipping AI-assisted anything at scale, the question isn't whether you can produce more. You already can. The question is whether you've built the thing that tells the gold from the slag before it lands on someone else's desk.
Live your dream — just don't make someone else clean up after it.
— Wayne