Graphic of AI robots
|

A New Tool to Stop AI From Stealing Your Content

If you’ve been wondering whether there’s anything you can actually do about AI systems scraping your blog content without permission, you’re not alone. For a long time, the answer has frustratingly been: not much.

But that might be starting to change.

The Really Simple Licensing (RSL) Standard is a new protocol that lets website owners add licensing terms for bots that crawl their content to train AI systems. Here’s what the RSL Standard is, how it works, and what it could mean for your website.

Legal Disclaimer: This post is for educational purposes only and does not constitute legal advice. Read full disclaimers.

What Is the RSL Standard?

The RSL Standard (short for Really Simple Licensing) is a new open protocol that lets website owners attach licensing terms to their content for AI crawlers. Instead of simply allowing or blocking bots through your website’s robots.txt, the RSL Standard adds a way to tell AI companies: “You can use this content, but only under specific licensing terms.”

It’s designed to help both large and small publishers signal that their content isn’t free training data for commercial AI models. While it doesn’t yet guarantee compliance or payment, it lays the groundwork for future enforcement and monetization.


Why the RSL Standard Matters to Bloggers

If you’ve ever published a blog post only to later find your words, structure, or even phrasing seemingly echoed by an AI-generated summary, you’re not imagining things.

AI companies have been training their models on massive amounts of content pulled from across the open web. And while major publishers who have leverage have started cutting licensing deals with companies like OpenAI, individual creators have been left with two bad options: ignore it or accept it.

The RSL Standard is designed to change that equation. It gives website owners a standardized way to implement licensing terms applicable to AI bots.


How It Works (In Plain English)

At its core, the RSL Standard builds on something that already exists: the robots.txt file on your website. That’s the file that tells bots (like Google’s crawler or ChatGPT’s GPTBot) which parts of your site they’re allowed to access. Historically, it’s been a binary setup: you either allow or disallow access.

The RSL Standard adds a third option. Instead of just yes or no, you can say: “Yes, you may crawl this, but here are the terms, including licensing, attribution, and/or royalty expectations.”

The standard was developed by the RSL Collective, and it has backing from big publishers like Reddit, Yahoo, Quora, and Medium, platforms that have every incentive to push back on unlicensed AI use. But here’s the catch: none of this works unless AI companies choose to honor it. Right now, compliance is voluntary, and some bots already ignore robots.txt entirely. For this to become more than a polite suggestion, pressure will need to come from publishers, collectives, and potentially regulators.


What the RSL Standard Can and Can’t Do (Right Now)

The idea behind the RSL Standard is straightforward: if your content is valuable enough to train commercial AI products, it should be treated like intellectual property, not public domain. That means AI companies should either license it or leave it alone.

That’s the vision. But where does that leave us right now?

Want to save this page?

I'll email this page to you, so you can come back to it later!

To learn how we protect your data see our privacy policy (link in footer).

Here’s what the RSL Standard can do:

  • Put your expectations on the record. It gives bloggers a clear, standardized way to say content isn’t free for AI to take. Even if it’s not enforceable yet, it’s a formal statement of intent.
  • Support collective action. Through the RSL Collective, publishers can push back together. It’s harder for AI companies to ignore a coordinated block than a single site.
  • Lay groundwork for monetization. If AI companies start licensing content (voluntarily or under pressure), bloggers who’ve already adopted RSL may be better positioned to benefit sooner.

Here’s what it can’t do (yet):

  • Force compliance. There’s no technical enforcement. Many bots already ignore robots.txt entirely.
  • Create a legally binding agreement. Posting restrictions on AI scraping doesn’t equal a contract, it’s more like a public notice.
  • Guarantee income. There’s no licensing revenue built in (yet). This is about future positioning.

In short, the RSL Standard is currently more of a marker than a shield. But in a fast-moving digital and legal landscape, planting a marker early can be a smart move.


What Bloggers Can Do Right Now

You don’t need to jump into code or change your setup overnight. But if you want to be proactive, here are a few smart steps to take now:

1. Join the RSL Collective

The RSL Collective is the group behind the licensing standard, and membership is open to individual bloggers, not just major media outlets. Joining gives you access to updates, implementation tools, and a voice in the broader movement.
Join the Collective

2. Add a clear statement about your AI content policy

Use your site’s footer, terms page, or about page to explain how you expect your content to be treated by AI systems. You might indicate that AI training use is not permitted, or that licensing is required. It won’t stop scraping, but it helps document your intent if licensing standards or enforcement mechanisms develop later.

3. If you’re comfortable with code, you can implement RSL terms now

The RSL Collective has published sample robots.txt code that lets you specify licensing expectations for AI crawlers whether that’s payment or attribution. If you manage your own site or work with a developer, you can add it now. → See the current implementation guide

And if you’d rather go the plugin route, here’s a new beta option for WordPress users.

4. If you use Cloudflare, keep an eye on Pay-Per-Crawl

While it’s not affiliated with the RSL Collective, it’s worth mentioning that Cloudflare is testing a tool that allows publishers to charge AI bots for accessing content, which is currently in private beta testing. If you’re already on Cloudflare, you’ll be among the first to access it when it’s publicly available.


RSL Is Still in Early Stages, But It’s Worth Paying Attention To

Right now, unfortunately, the RSL Standard can’t stop AI companies from scraping your blog. And while it won’t generate instant revenue or enforce your licensing terms in court, it does offer a structured way to signal that your content isn’t up for unrestricted use.

That may not seem like much, but if AI licensing frameworks develop further (and they will), having those signals in place now puts you in a better position to respond or participate.

This is early-stage infrastructure. It may evolve. It may stall. But it’s one of the few serious efforts aimed at making sure content creators aren’t left out of the AI economy.