Every domain here was asked for its homepage once as a browser and once as each of the crawlers behind ChatGPT, Claude, Perplexity, Gemini, Meta AI, Apple Intelligence and Doubao, and the answers compared. Not a prediction, not a crawl of someone else's dataset: the requests are made, and what came back is what you see. Google-Extended and Applebot-Extended never crawl — they are robots.txt opt-out tokens — so for those only robots.txt is reported, and a dash means there is nothing a request could have measured.
No crawler-level change recorded yet. A change is only reportable between two passes that both carry a per-crawler fingerprint, and the index started storing those on 2026-09-19: the score history before that says a site moved, not which crawler moved. This is an empty result, not a claim that nothing changed — and nobody can reconstruct it afterwards, which is why it is recorded from now on.
| Crawler | Measured | Served | robots.txt says no | Server says no anyway |
|---|---|---|---|---|
| ClaudeBot (Claude) | 47,615 | 84% | 2,084 | 6,651 |
| GPTBot (ChatGPT) | 47,615 | 85% | 2,604 | 6,332 |
| OAI-SearchBot (ChatGPT Search) | 47,615 | 90% | 828 | 4,562 |
| PerplexityBot (Perplexity) | 47,615 | 90% | 1,236 | 4,445 |
| Google-Extended (Gemini, AI Overviews) a robots.txt token, not a crawler: it controls how already-crawled pages may be used, and never makes a request of its own | 47,615 | — | 1,780 | — |
| Meta-ExternalAgent (Meta AI) | 22,120 | 85% | 848 | 3,096 |
| Amazonbot (Alexa, Rufus) | 22,120 | 81% | 928 | 3,926 |
| Bytespider (Doubao, Lark) | 3,830 | 88% | 64 | 427 |
| Applebot (Siri, Apple Intelligence) | 3,830 | 96% | 9 | 164 |
| Applebot-Extended (Apple Intelligence training) a robots.txt token, not a crawler: it controls how already-crawled pages may be used, and never makes a request of its own | 3,830 | — | 39 | — |
The last column is the number that exists nowhere else: robots.txt lets the crawler in and the server refuses it regardless. It is almost never a decision anyone made — it is an edge rule nobody checked.
| Edge in front of the site | Domains | Requests robots.txt allows | Refused anyway |
|---|---|---|---|
| Akamai | 749 | 3,423 | 37% |
| 788 | 3,811 | 35% | |
| Sucuri | 62 | 333 | 21% |
| AWS CloudFront | 3,126 | 15,097 | 17% |
| Cloudflare | 26,708 | 130,790 | 13% |
| no known edge | 11,904 | 59,860 | 11% |
| DDoS-Guard | 297 | 1,538 | 10% |
| Azure Front Door | 309 | 1,491 | 9% |
| Fastly | 1,467 | 6,684 | 9% |
| Varnish | 406 | 1,862 | 8% |
| Alibaba | 92 | 485 | 7% |
| Qrator | 182 | 916 | 6% |
| Vercel | 577 | 2,997 | 5% |
| BunnyCDN | 109 | 555 | 5% |
| Imperva | 164 | 847 | 2% |
| Netlify | 198 | 1,050 | 2% |
Unit: one domain × one crawler. Counted only where that site's robots.txt allows that crawler, so every refusal here contradicts the site's own stated policy. The edge is read from the response headers of the same request (cf-ray, akamai-grn, x-amz-cf-id, x-fastly-request-id…); sites with no recognisable signature are grouped as "no known edge", and domains measured before the header was recorded are left out of this table entirely rather than guessed into it. Read "no known edge" as an upper bound, not as a vendor: headers are not kept, so a domain read before a signature was added to the table stays in that bucket until it is re-requested — a sample of 250 of them re-requested on 19 Sep 2026 found 14% already carrying a signature the current table recognises. A weekly pass re-reads them, so the bucket shrinks on its own; the named vendors below are therefore undercounts, never overcounts. For part of the domains the edge was read in a later pass than the crawler verdicts (one request, headers only), so a site that changed CDN in between is shown under its current one until a full re-measurement replaces both. The two robots.txt-only tokens make no requests and are excluded.
| Edge | ClaudeBot | GPTBot | OAI-SearchBot | PerplexityBot | Meta-ExternalAgent | Amazonbot | Bytespider | Applebot |
|---|---|---|---|---|---|---|---|---|
| Akamai | 37% 2.97× | 38% 3.35× | 34% 4.12× | 37% 4.97× | 37% 3.21× | 40% 2.98× | — | — |
| 38% 3.06× | 38% 3.34× | 36% 4.31× | 35% 4.72× | 32% 2.79× | 30% 2.23× | — | — | |
| Sucuri | 36% 2.89× | 30% 2.65× | 31% 3.73× | 2% 0.22× | — | — | — | — |
| AWS CloudFront | 17% 1.34× | 18% 1.62× | 15% 1.83× | 15% 1.99× | 17% 1.50× | 18% 1.34× | 19% 0.86× | 7% 1.07× |
| Cloudflare | 15% 1.18× | 14% 1.25× | 9% 1.02× | 9% 1.20× | 16% 1.37× | 23% 1.75× | 9% 0.41× | 4% 0.54× |
| no known edge baseline | 12% | 11% | 8% | 7% | 12% | 13% | 23% | 7% |
| DDoS-Guard | 11% 0.87× | 12% 1.09× | 9% 1.10× | 7% 0.96× | 10% 0.90× | 12% 0.88× | — | — |
| Azure Front Door | 10% 0.79× | 9% 0.82× | 9% 1.05× | 8% 1.09× | 11% 0.98× | 6% 0.44× | — | — |
| Fastly | 11% 0.86× | 10% 0.92× | 7% 0.80× | 6% 0.83× | 9% 0.79× | 9% 0.66× | 10% 0.46× | 4% 0.60× |
| Varnish | 10% 0.78× | 10% 0.88× | 6% 0.66× | 6% 0.75× | 9% 0.76× | 10% 0.73× | — | — |
| Alibaba | 5% 0.43× | 8% 0.68× | 5% 0.65× | 7% 0.88× | 7% 0.59× | 9% 0.65× | — | — |
| Qrator | 4% 0.36× | 10% 0.93× | 7% 0.86× | 3% 0.45× | 7% 0.64× | 5% 0.40× | — | — |
| Vercel | 5% 0.44× | 6% 0.50× | 5% 0.54× | 5% 0.66× | 5% 0.45× | 6% 0.43× | — | — |
| BunnyCDN | 5% 0.38× | 5% 0.43× | 6% 0.67× | 3% 0.38× | 8% 0.65× | 6% 0.45× | — | — |
| Imperva | 3% 0.25× | 2% 0.17× | 1% 0.15× | 1% 0.08× | 3% 0.25× | 4% 0.30× | — | — |
| Netlify | 3% 0.20× | 3% 0.23× | 3% 0.30× | 3% 0.34× | 1% 0.07× | 1% 0.06× | — | — |
Each cell: of the domain × crawler pairs behind that vendor whose robots.txt allows that crawler, the share the server refused anyway. A vendor that refuses every crawler at the same rate is a wall nobody aimed; a vendor whose rate swings between crawlers is a managed list that names some user-agents and not others — which is the vendor's policy, not the site's. The small figure under each rate is that cell divided by the same cell for domains with no known edge: sites refuse AI crawlers for their own reasons everywhere, and this ratio is what being behind that vendor adds. 1.00× means the vendor changes nothing for that crawler. On Cloudflare (26,708 domains), the widest such gap is Amazonbot, refused on 23% of its 9,818 allowed pairs, against Applebot at 4% of 3,086 — 6.4×. Cells with fewer than 50 robots-allowed pairs are left empty rather than estimated; a crawler most sites block in robots.txt (Bytespider) reaches that floor on the largest vendors only. Columns do not share a denominator: an agent added to the registry later has only been asked on the domains measured since, so its column is a more recent slice of the same list — the pair count behind every cell is in its tooltip, and comparing two columns compares two samples, not two moments of one.
| Platform | Stores | Mean score | Open to all |
|---|---|---|---|
| shopify | 14,177 | 72.1 | 99% |
| woocommerce | 2,016 | 83.4 | 75% |
| other | 744 | 86.1 | 71% |
| bigcommerce | 445 | 69.8 | 96% |
| magento | 236 | 74.2 | 33% |
https://shop.lumnika.com/ai-readiness/mcp (no key, no signup) and ask it whether a domain lets AI
crawlers in, for the aggregate state of the web, or for the per-vendor edge-blocking table. It also measures
any domain on demand (measure_domain): 9 real requests made while you wait, even for sites the
index has not reached yet — and what it measures for you enters the public index on the next pass.
https://shop.lumnika.com/ai-readiness/api/v1/domain_readiness?host=example.com. The list of endpoints is at
/ai-readiness/api/v1 and the machine-readable contract at
openapi.json (OpenAPI 3.1) — both generated from the same tool
registry the MCP server publishes, so they cannot drift apart from what the index actually answers.
bot_* column per agent. Crawler columns carry what the server did (served,
blocked, no-answer…); the two robots.txt-only tokens carry
robots-allowed / robots-blocked, because no request is ever made for them. ·
Charts and series: the live index. Updated continuously.