Engraved alchemical cover artwork for “The Copyright Question Every AI Content Tool Has to Answer”

The Copyright Question Every AI Content Tool Has to Answer

I read TechCrunch's piece on the legality of training AI models on copyrighted books with the particular attention of someone who's built tools that sit right next to this question, not three steps removed from it. BookMasher generates book content. Article2Video turns written material into video. Both live downstream of the same models everyone's now arguing about in court.

The honest answer, as the article makes clear, is that nobody fully knows yet. Fair use in the US is being tested case by case, the UK's stance is still forming, and the big labs have mostly taken the "ask forgiveness not permission" approach with training data. That's not a comfortable place to build a business on, but it's the place we're all building on, whether we admit it or not.

Skin in the game

I don't train foundation models. I build on top of them. But that doesn't put me outside the argument — it puts me squarely inside it, just one layer removed. If the models underneath my tools got their training data through legally dubious means, that's not abstract to me. It's a supply chain question, the same as asking where your timber came from before you build furniture with it.

Forty years in software has taught me that "it's someone else's problem" is rarely true for long. When the licensing model for a platform you depend on changes, or a court rules against the way it was built, that cost lands on everyone downstream eventually. Anyone running content-automation tools who isn't thinking about this is postponing a conversation, not avoiding it.

What "complicated" actually means for builders

The TechCrunch piece is right to call it complicated, but complicated isn't the same as unknowable. A few things are becoming clearer even amid the mess:

Output matters more than input, legally speaking, for now. Generating content that's original in form — a summary, a restructured guide, a video built from your own script — sits on firmer ground than reproducing something close to a protected work. This is why I've always pushed the Masher tools toward transformation rather than replication. Take raw material, genuinely change its shape, add real structure and value. That's not just good practice for quality, it's the more defensible position if this all shakes out badly for the industry.

Provenance is becoming a selling point, not just a legal hedge. Users are starting to ask where content assistance tools get their capabilities from. I'd rather be able to answer that honestly than dodge the question. It's the same instinct that made me open about how the Masher suite works under the hood rather than dressing it up as some black box of magic.

Nobody's building on solid ground, so build for change. Whatever the courts eventually decide, the rules around AI training data are going to shift more than once in the next few years. Any tool architecture that assumes today's legal grey area is permanent is building on sand. I've deliberately kept the Masher tools flexible about which models and content sources they draw on, partly for quality reasons, but partly because I don't want to be locked into a legal position I didn't choose.

The part people don't say out loud

There's a temptation in this industry to treat copyright uncertainty as someone else's fight — the model providers', the publishers', the lawyers'. But if you're selling a tool that helps someone write a book, or turn an article into a video, you're part of the value chain that copyright law exists to police. Pretending otherwise doesn't make you safer, it just means you're not ready when the ground shifts.

My approach has always been the alchemist's one: raw material in, something genuinely transformed out. Not because it sounds nice on the website, but because transformation is both the better product and the more defensible one. If the input is murky, the least I can do is make sure what comes out the other end is unmistakably mine, and unmistakably useful in a way the original never was.

The legal question will get answered eventually, probably messily, probably in pieces, over several years and several jurisdictions. In the meantime, the builders who come out the other side in decent shape will be the ones who treated "complicated" as a design constraint, not an excuse to look away.

— Wayne