Can AI assistants read openbeautyfacts.org?
Measured on 2026-09-19 by asking the site 7 times — once as a browser, once as each of the
6 crawlers that feed ChatGPT, Claude, Perplexity, Gemini, Meta AI, Apple Intelligence and Doubao — and comparing
what came back. The AI-training opt-out tokens that never crawl are read from robots.txt instead.
D51 / 100
0 of 6 crawlers can read this page.
More readable than 7% of the 47,315 sites measured so far.
Structured data: WebSite, Organization.
What each crawler got back
| Crawler | HTTP | Result | robots.txt |
|---|
| ClaudeBot (Claude) | 200 |
thin |
blocked |
| GPTBot (ChatGPT) | 200 |
thin |
blocked |
| OAI-SearchBot (ChatGPT Search) | 200 |
allowed by the server, blocked in robots.txt |
blocked |
| PerplexityBot (Perplexity) | 200 |
thin |
blocked |
| Google-Extended (Gemini, AI Overviews) a robots.txt token, not a crawler: it controls how already-crawled pages may be used, and never makes a request of its own | — |
not opted out |
allowed-by-star |
| Meta-ExternalAgent (Meta AI) | 200 |
thin |
allowed-by-star |
| Amazonbot (Alexa, Rufus) | 200 |
thin |
blocked |
What to change, in order
- Let in the 1 crawler your own robots.txt already allows+1 crawler
Meta-ExternalAgent is turned away before reading the page (answered, but with far less text than a browser gets), while robots.txt permits it — so this block is not written in your site. No CDN signature was found in the response headers, so the refusal comes from the origin server itself or from a WAF this index does not recognise. Fixing it takes this domain from 0 to 1 of 6 crawlers.
- Publish sitemap.xml and llms.txt
sitemap.xml and llms.txt are missing, so an assistant has to discover the site by following links. llms.txt is the emerging convention for telling an assistant which pages actually matter.
- Fix the plain structure: one <h1>, a title, a description, alt text
The structure check scores 29/100. These are the cheapest signals on the page and the first ones an assistant uses to decide what the site is.
- 5 crawlers are named and refused in robots.txt — a deliberate choice
ClaudeBot, GPTBot, OAI-SearchBot, PerplexityBot, Amazonbot are blocked by a rule that names them. Nothing to fix here: this page records what is true, not what should be. It is listed so the deliberate part of the block is not confused with the accidental part above.
The checks
- Answers AI agents like it answers people29%
- Readable without JavaScript100%
- robots.txt lets the crawlers in29%
- Facts in JSON-LD84%
- Content reachable, not buried60%
- Publishes a map of itself0%
- Plain structure29%
Re-run this audit live →
Full AI-visibility report for openbeautyfacts.org →
Browse the whole index →
Embed this score
Put the badge on openbeautyfacts.org — it links back here, and re-measures every time this index re-crawls.
<a href="https://shop.lumnika.com/ai-readiness/openbeautyfacts.org"><img src="https://shop.lumnika.com/ai-readiness/openbeautyfacts.org/badge.svg" alt="AI readability"></a>