CrawlerToll

RSL licence lines

RSL — Really Simple Licensing — is the robots.txt-extension vocabulary for declaring AI-licensing terms. The v1.0 specification was published 2025-12-10 by the RSL Technical Steering Committee. Coalition members include Reddit, Yahoo, People Inc., Medium, Quora, O'Reilly, Ziff Davis, Stack Overflow, Fastly, and Cloudflare.

What RSL defines, and what CrawlerToll adds

RSL 1.0 defines a License: directive in robots.txt that points at an RSL XML licence file. That is the part that comes from the standard.

The Permits, Prohibits and Compensation lines are CrawlerToll's own convention. They are written in the same style, so they are called RSL-style licence lines here. CrawlerToll does not claim that a site using them conforms to RSL 1.0.

How CrawlerToll uses it

The WordPress plugin adds your rules to your site's own /robots.txt, under a comment line, whenever the site is public and enforcement is on. You edit them under Settings → CrawlerToll, in "Your AI-crawler rules (advanced)"; the defaults are sensible and you normally never touch them.

User-agent: GPTBot
Disallow: /
License: https://example.com/.well-known/context-license.json
Permits: ai-search, rag
Prohibits: ai-training, redistribution-without-attribution
Compensation: per-crawl 5000 micros USD
  • The License line points at the licence file the plugin serves on your site, /.well-known/context-license.json. See Context License.
  • The Compensation line follows the price you set. It is left out while the site cannot be paid, so no price is advertised that nobody could pay.
  • The same lines drive the plugin's decision for a declared AI crawler. See the decision tree.

Directives

| Directive | Where it comes from | What | |---|---|---| | User-agent: <name> | Yes (1994) | Selects which agents the rules apply to | | Disallow: <path> | Yes (1994) | Block by default | | Allow: <path> | Yes (RFC 9309) | Open carve-out (longest-match wins) | | Crawl-delay: <n> | De facto | Seconds between requests | | Sitemap: <url> | Yes (sitemaps.org) | Sitemap location | | License: <url> | RSL 1.0 | Points at the licence (an RSL XML licence file) | | Permits: <use, use> | CrawlerToll convention | Machine-readable permitted uses | | Prohibits: <use, use> | CrawlerToll convention | Machine-readable prohibited uses | | Compensation: <model> <price> micros <currency> [<url>] | CrawlerToll convention | Triggers 402 on blocked paths |

Use vocabulary

CrawlerToll uses one vocabulary for both Permits and Prohibits:

  • ai-training — bulk-corpus training data
  • ai-search — search-index style retrieval
  • ai-inference — live agent retrieval at inference time
  • rag — retrieval-augmented generation
  • agent-task — autonomous agent task completion
  • evaluation — benchmarks, eval sets
  • research — academic / non-commercial research
  • commercial-use, non-commercial-use
  • redistribution-with-attribution, redistribution-without-attribution
  • rebadging — claiming the content as your own
  • third-party-resale
  • competitive-dataset-creation
  • training-without-license

Compensation models

The models CrawlerToll understands in a Compensation line:

  • free — no payment required
  • negotiate — contact for terms
  • subscription — flat monthly with no per-call metering
  • per-crawl <micros> <currency> — pay per request
  • per-token <micros> <currency> — pay per output token
  • per-document <micros> <currency> — pay per document retrieved

Matching precedence

CrawlerToll matches paths the way robots.txt does, with the RFC 9309 (2022) clarification: longest-match wins, Allow ties beat Disallow. So:

User-agent: GPTBot
Allow: /articles
Disallow: /articles

/articles/123 → allowed (tie at length 9, Allow wins).

User-agent: GPTBot
Allow: /
Disallow: /private

/private/x → disallowed (longer Disallow match wins). /public → allowed (no Disallow matches).

See also

  • Decision tree — how matchAgent + matchPath compose into a 402-or-allow verdict
  • HTTP 402 standard — what the 402 response looks like when Compensation is declared