Agent Web Index

How much of the web can AI assistants actually read?

Every domain here was asked for its homepage once as a browser and once as each of the crawlers behind ChatGPT, Claude, Perplexity, Gemini, Meta AI, Apple Intelligence and Doubao, and the answers compared. Not a prediction, not a crawl of someone else's dataset: the requests are made, and what came back is what you see. Google-Extended and Applebot-Extended never crawl — they are robots.txt opt-out tokens — so for those only robots.txt is reported, and a dash means there is nothing a request could have measured.

46,907domains measured
22%blocked to ≥1 assistant
78%open to all of them
73mean readability /100

Per crawler

CrawlerMeasuredServedrobots.txt says noServer says no anyway
ClaudeBot (Claude)46,907 84%2,0666,503
GPTBot (ChatGPT)46,907 85%2,5836,190
OAI-SearchBot (ChatGPT Search)46,907 90%8234,490
PerplexityBot (Perplexity)46,907 90%1,2294,366
Google-Extended (Gemini, AI Overviews)
a robots.txt token, not a crawler: it controls how already-crawled pages may be used, and never makes a request of its own
46,907 1,765
Meta-ExternalAgent (Meta AI)18,535 83%8262,931
Amazonbot (Alexa, Rufus)18,535 79%8973,652
Bytespider (Doubao, Lark)245 74%2554
Applebot (Siri, Apple Intelligence)245 88%530
Applebot-Extended (Apple Intelligence training)
a robots.txt token, not a crawler: it controls how already-crawled pages may be used, and never makes a request of its own
245 16

The last column is the number that exists nowhere else: robots.txt lets the crawler in and the server refuses it regardless. It is almost never a decision anyone made — it is an edge rule nobody checked.

Who is actually doing the blocking (8,608 domains with the edge identified)

Edge in front of the siteDomainsRequests robots.txt allowsRefused anyway
Google391 1,75160%
Akamai122 51336%
Cloudflare3,197 12,91021%
AWS CloudFront950 3,88016%
Azure Front Door58 24015%
no known edge2,952 12,02311%
Fastly452 1,7477%
Vercel133 5846%
Qrator61 2835%
Varnish64 2695%
DDoS-Guard63 2875%
Netlify67 3024%

Unit: one domain × one crawler. Counted only where that site's robots.txt allows that crawler, so every refusal here contradicts the site's own stated policy. The edge is read from the response headers of the same request (cf-ray, akamai-grn, x-amz-cf-id, x-fastly-request-id…); sites with no recognisable signature are grouped as "no known edge", and domains measured before the header was recorded are left out of this table entirely rather than guessed into it. For part of the domains the edge was read in a later pass than the crawler verdicts (one request, headers only), so a site that changed CDN in between is shown under its current one until a full re-measurement replaces both. The two robots.txt-only tokens make no requests and are excluded.

By shop platform (16,911 confirmed stores)

PlatformStoresMean scoreOpen to all
shopify14,010 7299%
woocommerce1,690 83.780%
other626 86.273%
bigcommerce435 69.696%
magento149 74.936%

Most readable

outbackaccounting.com.auA+ 100cmux.devA+ 100spotmedia.roA+ 100cambiocolombia.comA+ 100press.lvA+ 100haberler.comA+ 99colorlib.comA+ 99dynadot.comA+ 99creativethemes.comA+ 99reason.comA+ 99gong.ioA+ 99vedayuonline.comA+ 99canaltech.com.brA+ 99fogaonet.comA+ 99adminforge.deA+ 99skymesh.net.auA+ 99indoleads.comA+ 991filmyfly.onlineA+ 99dash0.comA+ 99thesouthafrican.comA+ 99getgrav.orgA+ 99leadinfo.comA+ 99server-eye.deA+ 99one.nzA+ 99gcaptain.comA+ 99

Least readable

f95zone.toF 14cargocollective.comF 14moneyconverterapp.comF 14lancasteronline.comF 15ctee.com.twF 16capitalizemytitle.comF 16ft.comF 17androidheadlines.comF 17ecourtsindia.comF 17winknews.comF 17milanuncios.comF 18ylilauta.orgF 18gogoroyal.comF 19alltheweb.comF 19watchpeopledie.tvF 20dmitory.comF 20cidadeverde.comF 20f6.securityF 21infojobs.netF 22rhizome.orgF 22ebay.usF 22printemps.comF 22placedestendances.comF 22bloomberg.comF 23t2.ruF 23
Audit your own site live →
Method and limits. One vantage point (Europe), one page per domain (the homepage), 12-second timeout per request. Blocking AI crawlers is a legitimate choice, not a failure: these pages record what is true, not what should be. 22,477 domains failed to answer and are excluded from every percentage above; 16,227 are infrastructure rather than sites and are counted apart. Rank-bearing domains come from the Tranco research list.

Full dataset: dataset.csv · Charts and series: the live index. Updated continuously.