Agent Web Index

How much of the web can AI assistants actually read?

Every domain here was asked for its homepage once as a browser and once as each of the crawlers behind ChatGPT, Claude, Perplexity, Gemini, Meta AI, Apple Intelligence and Doubao, and the answers compared. Not a prediction, not a crawl of someone else's dataset: the requests are made, and what came back is what you see. Google-Extended and Applebot-Extended never crawl — they are robots.txt opt-out tokens — so for those only robots.txt is reported, and a dash means there is nothing a request could have measured.

47,615domains measured
23%blocked to ≥1 assistant
77%open to all of them
73.1mean readability /100

Who changed their mind — the last 30 days

No crawler-level change recorded yet. A change is only reportable between two passes that both carry a per-crawler fingerprint, and the index started storing those on 2026-09-19: the score history before that says a site moved, not which crawler moved. This is an empty result, not a claim that nothing changed — and nobody can reconstruct it afterwards, which is why it is recorded from now on.

Per crawler

CrawlerMeasuredServedrobots.txt says noServer says no anyway
ClaudeBot (Claude)47,615 84%2,0846,651
GPTBot (ChatGPT)47,615 85%2,6046,332
OAI-SearchBot (ChatGPT Search)47,615 90%8284,562
PerplexityBot (Perplexity)47,615 90%1,2364,445
Google-Extended (Gemini, AI Overviews)
a robots.txt token, not a crawler: it controls how already-crawled pages may be used, and never makes a request of its own
47,615 1,780
Meta-ExternalAgent (Meta AI)22,120 85%8483,096
Amazonbot (Alexa, Rufus)22,120 81%9283,926
Bytespider (Doubao, Lark)3,830 88%64427
Applebot (Siri, Apple Intelligence)3,830 96%9164
Applebot-Extended (Apple Intelligence training)
a robots.txt token, not a crawler: it controls how already-crawled pages may be used, and never makes a request of its own
3,830 39

The last column is the number that exists nowhere else: robots.txt lets the crawler in and the server refuses it regardless. It is almost never a decision anyone made — it is an edge rule nobody checked.

Who is actually doing the blocking (47,174 domains with the edge identified)

Edge in front of the siteDomainsRequests robots.txt allowsRefused anyway
Akamai749 3,42337%
Google788 3,81135%
Sucuri62 33321%
AWS CloudFront3,126 15,09717%
Cloudflare26,708 130,79013%
no known edge11,904 59,86011%
DDoS-Guard297 1,53810%
Azure Front Door309 1,4919%
Fastly1,467 6,6849%
Varnish406 1,8628%
Alibaba92 4857%
Qrator182 9166%
Vercel577 2,9975%
BunnyCDN109 5555%
Imperva164 8472%
Netlify198 1,0502%

Unit: one domain × one crawler. Counted only where that site's robots.txt allows that crawler, so every refusal here contradicts the site's own stated policy. The edge is read from the response headers of the same request (cf-ray, akamai-grn, x-amz-cf-id, x-fastly-request-id…); sites with no recognisable signature are grouped as "no known edge", and domains measured before the header was recorded are left out of this table entirely rather than guessed into it. Read "no known edge" as an upper bound, not as a vendor: headers are not kept, so a domain read before a signature was added to the table stays in that bucket until it is re-requested — a sample of 250 of them re-requested on 19 Sep 2026 found 14% already carrying a signature the current table recognises. A weekly pass re-reads them, so the bucket shrinks on its own; the named vendors below are therefore undercounts, never overcounts. For part of the domains the edge was read in a later pass than the crawler verdicts (one request, headers only), so a site that changed CDN in between is shown under its current one until a full re-measurement replaces both. The two robots.txt-only tokens make no requests and are excluded.

Same vendors, one column per crawler

EdgeClaudeBotGPTBotOAI-SearchBotPerplexityBotMeta-ExternalAgentAmazonbotBytespiderApplebot
Akamai37%
2.97×
38%
3.35×
34%
4.12×
37%
4.97×
37%
3.21×
40%
2.98×
Google38%
3.06×
38%
3.34×
36%
4.31×
35%
4.72×
32%
2.79×
30%
2.23×
Sucuri36%
2.89×
30%
2.65×
31%
3.73×
2%
0.22×
AWS CloudFront17%
1.34×
18%
1.62×
15%
1.83×
15%
1.99×
17%
1.50×
18%
1.34×
19%
0.86×
7%
1.07×
Cloudflare15%
1.18×
14%
1.25×
9%
1.02×
9%
1.20×
16%
1.37×
23%
1.75×
9%
0.41×
4%
0.54×
no known edge
baseline
12%11%8%7%12%13%23%7%
DDoS-Guard11%
0.87×
12%
1.09×
9%
1.10×
7%
0.96×
10%
0.90×
12%
0.88×
Azure Front Door10%
0.79×
9%
0.82×
9%
1.05×
8%
1.09×
11%
0.98×
6%
0.44×
Fastly11%
0.86×
10%
0.92×
7%
0.80×
6%
0.83×
9%
0.79×
9%
0.66×
10%
0.46×
4%
0.60×
Varnish10%
0.78×
10%
0.88×
6%
0.66×
6%
0.75×
9%
0.76×
10%
0.73×
Alibaba5%
0.43×
8%
0.68×
5%
0.65×
7%
0.88×
7%
0.59×
9%
0.65×
Qrator4%
0.36×
10%
0.93×
7%
0.86×
3%
0.45×
7%
0.64×
5%
0.40×
Vercel5%
0.44×
6%
0.50×
5%
0.54×
5%
0.66×
5%
0.45×
6%
0.43×
BunnyCDN5%
0.38×
5%
0.43×
6%
0.67×
3%
0.38×
8%
0.65×
6%
0.45×
Imperva3%
0.25×
2%
0.17×
1%
0.15×
1%
0.08×
3%
0.25×
4%
0.30×
Netlify3%
0.20×
3%
0.23×
3%
0.30×
3%
0.34×
1%
0.07×
1%
0.06×

Each cell: of the domain × crawler pairs behind that vendor whose robots.txt allows that crawler, the share the server refused anyway. A vendor that refuses every crawler at the same rate is a wall nobody aimed; a vendor whose rate swings between crawlers is a managed list that names some user-agents and not others — which is the vendor's policy, not the site's. The small figure under each rate is that cell divided by the same cell for domains with no known edge: sites refuse AI crawlers for their own reasons everywhere, and this ratio is what being behind that vendor adds. 1.00× means the vendor changes nothing for that crawler. On Cloudflare (26,708 domains), the widest such gap is Amazonbot, refused on 23% of its 9,818 allowed pairs, against Applebot at 4% of 3,086 — 6.4×. Cells with fewer than 50 robots-allowed pairs are left empty rather than estimated; a crawler most sites block in robots.txt (Bytespider) reaches that floor on the largest vendors only. Columns do not share a denominator: an agent added to the registry later has only been asked on the domains measured since, so its column is a more recent slice of the same list — the pair count behind every cell is in its tooltip, and comparing two columns compares two samples, not two moments of one.

By shop platform (17,619 confirmed stores)

PlatformStoresMean scoreOpen to all
shopify14,177 72.199%
woocommerce2,016 83.475%
other744 86.171%
bigcommerce445 69.896%
magento236 74.233%

Most readable

outbackaccounting.com.auA+ 100cmux.devA+ 100spotmedia.roA+ 100cambiocolombia.comA+ 100press.lvA+ 100haberler.comA+ 99colorlib.comA+ 99dynadot.comA+ 99creativethemes.comA+ 99reason.comA+ 99gong.ioA+ 99vedayuonline.comA+ 99canaltech.com.brA+ 99fogaonet.comA+ 99adminforge.deA+ 99skymesh.net.auA+ 99indoleads.comA+ 991filmyfly.onlineA+ 99dash0.comA+ 99thesouthafrican.comA+ 99getgrav.orgA+ 99leadinfo.comA+ 99server-eye.deA+ 99one.nzA+ 99gcaptain.comA+ 99

Least readable

f95zone.toF 14cargocollective.comF 14moneyconverterapp.comF 14lancasteronline.comF 15ctee.com.twF 16capitalizemytitle.comF 16ft.comF 17androidheadlines.comF 17ecourtsindia.comF 17winknews.comF 17milanuncios.comF 18ylilauta.orgF 18gogoroyal.comF 19alltheweb.comF 19watchpeopledie.tvF 20dmitory.comF 20cidadeverde.comF 20f6.securityF 21infojobs.netF 22rhizome.orgF 22ebay.usF 22printemps.comF 22placedestendances.comF 22bloomberg.comF 23t2.ruF 23
Audit your own site live →
Method and limits. One vantage point (Europe), one page per domain (the homepage), 12-second timeout per request. Blocking AI crawlers is a legitimate choice, not a failure: these pages record what is true, not what should be. 22,544 domains failed to answer and are excluded from every percentage above; 16,284 are infrastructure rather than sites and are counted apart. Rank-bearing domains come from the Tranco research list.

Call it from an assistant. The index is also an MCP server — connect https://shop.lumnika.com/ai-readiness/mcp (no key, no signup) and ask it whether a domain lets AI crawlers in, for the aggregate state of the web, or for the per-vendor edge-blocking table. It also measures any domain on demand (measure_domain): 9 real requests made while you wait, even for sites the index has not reached yet — and what it measures for you enters the public index on the next pass.

Or call it as a plain API. Every MCP tool is also one GET, no key and no signup: https://shop.lumnika.com/ai-readiness/api/v1/domain_readiness?host=example.com. The list of endpoints is at /ai-readiness/api/v1 and the machine-readable contract at openapi.json (OpenAPI 3.1) — both generated from the same tool registry the MCP server publishes, so they cannot drift apart from what the index actually answers.

Full dataset: dataset.csv — one row per domain, one bot_* column per agent. Crawler columns carry what the server did (served, blocked, no-answer…); the two robots.txt-only tokens carry robots-allowed / robots-blocked, because no request is ever made for them. · Charts and series: the live index. Updated continuously.

Take the data. CC BY 4.0, no signup, mirrored every day with an immutable daily snapshot of the aggregate — so the series survives us: Hugging Face · GitHub · Zenodo (DOI).