CrawlerToll

Decision tree

On every front-end request to an ordinary page, the CrawlerToll WordPress plugin makes one small, explicit decision. This page lists it in order. Sealed articles are handled separately: see the unlock service.

What is never checked

The plugin does nothing for these, so they always pass:

  • the WordPress admin, the REST API, AJAX, WP-Cron and XML-RPC
  • /robots.txt and everything under /.well-known/, so crawlers can always read your terms
  • sealed articles, which are served as the preview plus the encrypted body to everyone

The tree

Search-engine or social-preview crawler?  ─────────────────▶  PASS
(Googlebot, Bingbot, DuckDuckBot and so on)
 
Page marked noindex by your SEO plugin?  ──────────────────▶  PASS
 
User agent not in the 30-crawler catalogue?  ──────────────▶  PASS
(people and everything else)
 
Declared AI crawler:
  │
  ├── your policy has no rule for it  ─────────────────────▶  PASS
  │
  └── your policy has a rule for it
        │
        ├── path allowed  ─────────────────────────────────▶  PASS
        │
        └── path disallowed
              │
              ├── the rule has a Compensation line
              │     │
              │     ├── the site can be paid  ─────────────▶  402 with the price
              │     │
              │     └── the site cannot be paid yet  ──────▶  403 not licensed
              │
              └── no Compensation line  ───────────────────▶  403 forbidden

"The site can be paid" means you have set a USDC payout address, your own Stripe keys, or a payment URL. Until then the plugin never quotes a price nobody could pay. A declared AI crawler is told the content is not licensed, with a pointer to your robots.txt.

What the crawler gets

402 with the price. The response carries Crawler-Price and Crawler-Price-Rail headers, a Link header, Retry-After, and a JSON body with the offer. The price is advertised, not charged: paid access is sold per sealed article. See HTTP 402.

403 not licensed. A JSON body with error: "not_licensed" and the address of your robots.txt. No price is quoted.

403 forbidden. A JSON body with error: "forbidden" and the list of rules that fired.

Every answer also carries X-CrawlerToll-Action and, for a declared crawler, X-CrawlerToll-Bot-Name and X-CrawlerToll-Operator headers, so you can see what the plugin decided.

Edge cases

A crawler that signs its requests

The plugin does not verify Web Bot Auth signatures. It recognises a crawler by the user agent it declares, so a signed request is treated like any other. See Web Bot Auth.

Multiple User-agent lines

RSL-style licence lines inherit robots.txt's "consecutive UA lines form one group" rule. So:

User-agent: GPTBot
User-agent: ClaudeBot
Disallow: /

is one group matching both user agents. The most specific user-agent token wins (longest substring match); a literal * is the catch-all of last resort.

Allow vs Disallow ties

Per RFC 9309 (2022), longest-match wins, and Allow ties beat Disallow:

User-agent: GPTBot
Allow: /articles
Disallow: /articles

/articles/123 → allowed (tie at length 9, Allow wins).

See also