Agent Web Index

How much of the web can AI assistants actually read?

Every domain here was asked for its homepage once as a browser and once as each of the crawlers behind ChatGPT, Claude, Perplexity, Gemini, Meta AI, Apple Intelligence and Doubao, and the answers compared. Not a prediction, not a crawl of someone else's dataset: the requests are made, and what came back is what you see. Google-Extended and Applebot-Extended never crawl — they are robots.txt opt-out tokens — so for those only robots.txt is reported, and a dash means there is nothing a request could have measured.

47,309domains measured
23%blocked to ≥1 assistant
77%open to all of them
73.1mean readability /100

Who changed their mind — the last 30 days

No crawler-level change recorded yet. A change is only reportable between two passes that both carry a per-crawler fingerprint, and the index started storing those on 2026-09-19: the score history before that says a site moved, not which crawler moved. This is an empty result, not a claim that nothing changed — and nobody can reconstruct it afterwards, which is why it is recorded from now on.

Per crawler

CrawlerMeasuredServedrobots.txt says noServer says no anyway
ClaudeBot (Claude)47,309 84%2,0776,577
GPTBot (ChatGPT)47,309 85%2,5956,256
OAI-SearchBot (ChatGPT Search)47,309 90%8264,522
PerplexityBot (Perplexity)47,309 90%1,2334,402
Google-Extended (Gemini, AI Overviews)
a robots.txt token, not a crawler: it controls how already-crawled pages may be used, and never makes a request of its own
47,309 1,775
Meta-ExternalAgent (Meta AI)20,745 84%8413,012
Amazonbot (Alexa, Rufus)20,745 80%9193,801
Bytespider (Doubao, Lark)2,455 89%48259
Applebot (Siri, Apple Intelligence)2,455 96%7104
Applebot-Extended (Apple Intelligence training)
a robots.txt token, not a crawler: it controls how already-crawled pages may be used, and never makes a request of its own
2,455 29

The last column is the number that exists nowhere else: robots.txt lets the crawler in and the server refuses it regardless. It is almost never a decision anyone made — it is an edge rule nobody checked.

Who is actually doing the blocking (46,866 domains with the edge identified)

Edge in front of the siteDomainsRequests robots.txt allowsRefused anyway
Akamai748 3,41137%
Google787 3,78835%
Sucuri61 32521%
AWS CloudFront3,114 14,99317%
Cloudflare26,541 125,66413%
no known edge11,807 58,77310%
DDoS-Guard296 1,53010%
Azure Front Door308 1,4839%
Fastly1,455 6,5398%
Varnish404 1,8468%
Alibaba91 4777%
Qrator181 9086%
Vercel570 2,9095%
BunnyCDN108 5475%
Imperva162 8312%
Netlify198 1,0462%

Unit: one domain × one crawler. Counted only where that site's robots.txt allows that crawler, so every refusal here contradicts the site's own stated policy. The edge is read from the response headers of the same request (cf-ray, akamai-grn, x-amz-cf-id, x-fastly-request-id…); sites with no recognisable signature are grouped as "no known edge", and domains measured before the header was recorded are left out of this table entirely rather than guessed into it. Read "no known edge" as an upper bound, not as a vendor: headers are not kept, so a domain read before a signature was added to the table stays in that bucket until it is re-requested — a sample of 250 of them re-requested on 19 Sep 2026 found 14% already carrying a signature the current table recognises. A weekly pass re-reads them, so the bucket shrinks on its own; the named vendors below are therefore undercounts, never overcounts. For part of the domains the edge was read in a later pass than the crawler verdicts (one request, headers only), so a site that changed CDN in between is shown under its current one until a full re-measurement replaces both. The two robots.txt-only tokens make no requests and are excluded.

Same vendors, one column per crawler

EdgeClaudeBotGPTBotOAI-SearchBotPerplexityBotMeta-ExternalAgentAmazonbotBytespiderApplebot
Akamai37%
2.99×
38%
3.37×
34%
4.15×
37%
5.01×
37%
3.21×
40%
2.99×
Google38%
3.08×
38%
3.36×
36%
4.34×
35%
4.75×
32%
2.81×
30%
2.27×
Sucuri35%
2.82×
29%
2.56×
30%
3.61×
2%
0.23×
AWS CloudFront17%
1.35×
18%
1.62×
15%
1.84×
15%
2.00×
18%
1.52×
18%
1.36×
8%
1.30×
Cloudflare15%
1.18×
14%
1.24×
9%
1.03×
9%
1.20×
17%
1.49×
25%
1.90×
8%
0.37×
4%
0.63×
no known edge
baseline
12%11%8%7%12%13%22%6%
DDoS-Guard11%
0.88×
12%
1.10×
9%
1.11×
7%
0.97×
10%
0.91×
12%
0.89×
Azure Front Door10%
0.80×
9%
0.82×
9%
1.06×
8%
1.11×
11%
0.99×
6%
0.45×
Fastly10%
0.83×
10%
0.87×
7%
0.79×
6%
0.84×
9%
0.76×
9%
0.66×
Varnish10%
0.78×
10%
0.89×
6%
0.67×
6%
0.76×
9%
0.77×
10%
0.74×
Alibaba5%
0.44×
8%
0.69×
5%
0.66×
7%
0.89×
7%
0.61×
9%
0.66×
Qrator4%
0.36×
11%
0.94×
7%
0.87×
3%
0.45×
7%
0.64×
5%
0.41×
Vercel6%
0.44×
6%
0.51×
5%
0.55×
5%
0.67×
5%
0.47×
6%
0.46×
BunnyCDN5%
0.39×
5%
0.44×
6%
0.68×
3%
0.38×
8%
0.66×
6%
0.46×
Imperva3%
0.26×
2%
0.17×
1%
0.15×
1%
0.09×
3%
0.26×
4%
0.31×
Netlify3%
0.21×
3%
0.23×
3%
0.31×
3%
0.34×
1%
0.07×
1%
0.06×

Each cell: of the domain × crawler pairs behind that vendor whose robots.txt allows that crawler, the share the server refused anyway. A vendor that refuses every crawler at the same rate is a wall nobody aimed; a vendor whose rate swings between crawlers is a managed list that names some user-agents and not others — which is the vendor's policy, not the site's. The small figure under each rate is that cell divided by the same cell for domains with no known edge: sites refuse AI crawlers for their own reasons everywhere, and this ratio is what being behind that vendor adds. 1.00× means the vendor changes nothing for that crawler. On Cloudflare (26,541 domains), the widest such gap is Amazonbot, refused on 25% of its 8,702 allowed pairs, against Applebot at 4% of 1,963 — 6.9×. Cells with fewer than 50 robots-allowed pairs are left empty rather than estimated; a crawler most sites block in robots.txt (Bytespider) reaches that floor on the largest vendors only. Columns do not share a denominator: an agent added to the registry later has only been asked on the domains measured since, so its column is a more recent slice of the same list — the pair count behind every cell is in its tooltip, and comparing two columns compares two samples, not two moments of one.

By shop platform (17,311 confirmed stores)

PlatformStoresMean scoreOpen to all
shopify14,109 72.199%
woocommerce1,863 83.677%
other685 86.372%
bigcommerce443 69.896%
magento210 74.534%

Most readable

outbackaccounting.com.auA+ 100cmux.devA+ 100spotmedia.roA+ 100cambiocolombia.comA+ 100press.lvA+ 100haberler.comA+ 99colorlib.comA+ 99dynadot.comA+ 99creativethemes.comA+ 99reason.comA+ 99gong.ioA+ 99vedayuonline.comA+ 99canaltech.com.brA+ 99fogaonet.comA+ 99adminforge.deA+ 99skymesh.net.auA+ 99indoleads.comA+ 991filmyfly.onlineA+ 99dash0.comA+ 99thesouthafrican.comA+ 99getgrav.orgA+ 99leadinfo.comA+ 99server-eye.deA+ 99one.nzA+ 99gcaptain.comA+ 99

Least readable

f95zone.toF 14cargocollective.comF 14moneyconverterapp.comF 14lancasteronline.comF 15ctee.com.twF 16capitalizemytitle.comF 16ft.comF 17androidheadlines.comF 17ecourtsindia.comF 17winknews.comF 17milanuncios.comF 18ylilauta.orgF 18gogoroyal.comF 19alltheweb.comF 19watchpeopledie.tvF 20dmitory.comF 20cidadeverde.comF 20f6.securityF 21infojobs.netF 22rhizome.orgF 22ebay.usF 22printemps.comF 22placedestendances.comF 22bloomberg.comF 23t2.ruF 23
Audit your own site live →
Method and limits. One vantage point (Europe), one page per domain (the homepage), 12-second timeout per request. Blocking AI crawlers is a legitimate choice, not a failure: these pages record what is true, not what should be. 22,498 domains failed to answer and are excluded from every percentage above; 16,248 are infrastructure rather than sites and are counted apart. Rank-bearing domains come from the Tranco research list.

Call it from an assistant. The index is also an MCP server — connect https://shop.lumnika.com/ai-readiness/mcp (no key, no signup) and ask it whether a domain lets AI crawlers in, for the aggregate state of the web, or for the per-vendor edge-blocking table. It also measures any domain on demand (measure_domain): 9 real requests made while you wait, even for sites the index has not reached yet — and what it measures for you enters the public index on the next pass.

Or call it as a plain API. Every MCP tool is also one GET, no key and no signup: https://shop.lumnika.com/ai-readiness/api/v1/domain_readiness?host=example.com. The list of endpoints is at /ai-readiness/api/v1 and the machine-readable contract at openapi.json (OpenAPI 3.1) — both generated from the same tool registry the MCP server publishes, so they cannot drift apart from what the index actually answers.

Full dataset: dataset.csv — one row per domain, one bot_* column per agent. Crawler columns carry what the server did (served, blocked, no-answer…); the two robots.txt-only tokens carry robots-allowed / robots-blocked, because no request is ever made for them. · Charts and series: the live index. Updated continuously.

Take the data. CC BY 4.0, no signup, mirrored every day with an immutable daily snapshot of the aggregate — so the series survives us: Hugging Face · GitHub · Zenodo (DOI).