If you run a small business, the honest first question about all of this is whether it's real or whether someone is selling you a new anxiety. So before we measured any customer, we measured the top of the market: 34 well-known companies, checked for exactly the things the GEO industry tells you to go and do. Here is everything we found, including the parts that argue against buying anything.
On 28 August 2026 we requested a fixed list of URLs from 34 websites and read what came back. Zero AI calls, zero cost, no interpretation in the middle — this is a file-and-header scan, the same mechanical checks that run against your own site in the audit section of a report. That is also the limit of it, and we come back to that below.
The 34, in three groups:
For each one we asked for 10 AI-facing files (llms.txt,
agents.md, ai.txt, an agentic-discovery sitemap,
.well-known/mcp.json and others), read
robots.txt and checked it against 22 named AI crawlers, and
pulled the homepage to count its structured data, its h1
elements and how many words of readable text the HTML contains before any
JavaScript runs.
Three of the 34 refused us. Warby Parker returned 403,
Patagonia 404 and away.com dropped the connection — bot mitigation
reacting to a datacentre IP address, almost certainly, not a statement
about their sites. Their homepages are therefore unmeasured, not zero, and
every count below is out of the 31 homepages and 30
robots.txt files we could actually read. Roborock's
robots.txt also failed to load.
llms.txt is the flagship deliverable of a lot of GEO
packages. Across the 34:
| File | Found on |
|---|---|
/llms.txt | 22 of 34 |
/llms-full.txt | 12 of 34 |
/agents.md (or /AGENTS.md) | 9 of 34 |
/sitemap_agentic_discovery.xml | 7 of 34 |
/.well-known/mcp.json | 4 of 34 |
/.well-known/ai-plugin.json | 1 of 34 |
/ai.txt | 0 of 34 |
/.well-known/agents.md | 0 of 34 |
/mcp | 0 of 34 |
It splits sharply by industry. 13 of the 16 software companies serve an
llms.txt; 4 of the 8 DTC brands and 5 of the 10 China-based
brands do. Companies whose customers are developers publish developer
files. That is a less exciting explanation than "the leaders are ahead of
you", and it fits the data better.
Two of these deserve saying out loud because they get sold anyway:
nothing in this sample served an ai.txt, and
Google's own documentation says it does not read llms.txt.
A file is not a switch.
Nine sites carry an agents.md — the file that tells an AI
agent who you are and how to deal with you. We read all nine.
Six of them are the same file.
Allbirds, Glossier, Gymshark, Brooklinen, EcoFlow and UGREEN each serve a
4,129–4,311 byte document headed
# Agent Instructions — {brand}, and each of the six points
an AI agent at shop.app — Shopify's own checkout — exactly
eight times. Compared character by character, Allbirds'
copy and UGREEN's differ only in the brand name and one policy link. Five
of the six also serve an llms.txt containing exactly five
links.
We have seen this file before. It was also on the site of a one-person handmade skincare business we scanned while building the method — same template, same eight pointers to someone else's checkout. So a large athleisure brand and a soap maker working out of a spare room currently introduce themselves to AI with the same paragraph, and neither of them wrote it.
Three sites wrote their own. Shopify's runs to 12,522 bytes, Zapier's to
1,205, and Anker's to 6,610 — Anker being the only brand in the sample
that took the platform default and rewrote it, dropping
shop.app from it entirely and adding sections on regional
markets and canonical product links.
The point for you is not that these files are worthless. It's that having one and having written one are different states, and any audit that reports the first as a green tick is measuring your hosting provider.
Where a company clearly did write its own llms.txt, the
dominant format is a plain annotated list of its own pages. Stripe's
carries 288 links, HubSpot's 201, SHEIN's 190. The platform-generated ones
carry five.
That is a description of what they did, not evidence that it worked. We did not measure whether any of these files changed how an assistant talks about the company, and we are not aware of anyone who has published that measurement either.
This is the question small-business owners ask us most: should I block the
AI bots? Of the 30 robots.txt files we could read, only
7 mention any of the 22 AI crawlers we checked, and only
3 block one:
| Site | AI crawlers fully disallowed |
|---|---|
| Canva | 12 |
| Figma | 7 |
| Notion | 2 |
| The other 27 readable sites | none |
The three who do block are all companies whose product is a design or document canvas, i.e. companies with an obvious training-data reason of their own. Everyone else lets them in. If someone has told you to shut the AI crawlers out as a defensive measure, the top of the market is doing the opposite.
This is the part that changed how we write reports. Out of the 31 homepages we could read:
| Check | Head-of-market sites passing |
|---|---|
At least one h1 on the homepage | 21 of 31 |
| Any JSON-LD structured data at all | 20 of 31 |
Organization schema | 14 of 31 |
AggregateRating schema | 2 of 31 |
Review schema | 0 of 31 |
Ten of these companies have no h1 on their homepage. Eleven
publish no structured data whatsoever. And review markup — a standing item
on essentially every SEO checklist ever written — appears on two sites out
of 31 in the form of an aggregate rating, and on none at all as individual
reviews.
Readable text is uneven too. The median homepage served 1,059 words of text in its HTML before any JavaScript ran; the largest served 40,745. Four served under 300: UGREEN 72 words, Gymshark and Narwal one word each, Temu none. Their pages plainly work for humans in a browser — but an AI system reading the raw HTML sees close to nothing, and we did not run JavaScript, so all this measures is what arrives in the first response.
Nearer than the marketing suggests, on the mechanical checks at least. The column on the right is the five small businesses we ran the full method against while validating it. We don't name them and we don't publish any individual result — that rule applies to everyone who buys a report, which is why the sample report is pseudonymised too. Five is a tiny sample and the numbers below should be read as a sketch, not a rate.
| Check | Head of market | Five small businesses |
|---|---|---|
h1 on the homepage | 21 of 31 | 1 of 5 |
Organization schema | 14 of 31 | 2 of 5 |
Review / AggregateRating schema | 2 of 31 | 0 of 5 |
robots.txt naming any AI crawler | 7 of 30 | 0 of 5 |
Read the rows in order. The first is a genuine gap and worth ten minutes of somebody's time. The second is a gap against a majority that isn't itself a majority. The third and fourth are not gaps at all — nobody is doing them, at any size, and a report that listed them as your failings would be padding.
That is now a rule in how we write: a check has to be compared against this benchmark before it can be called a problem. Below the benchmark is a problem. Level with it is at most an opportunity. Above it gets said out loud as a win. It removed two items from our audit list outright, and it is the reason your report will not be a 60-line sheet of things you have not done.
It is a file scan. It measures what these companies have deployed, and that is all it measures.
Three things, and none of them is "buy something".
llms.txt, two thirds of the ones that do had it written by
their shopping platform, and eleven of them publish no structured data
at all. Anyone quoting you thousands to catch up with the leaders should
be asked which leaders, and shown this.robots.txt, for free,
right now.The only way to know the second thing is to ask the assistants, in your category, and look at what comes back. That's the measurement, and you can start it without paying us.
23 questions across 3 AI surfaces, how a mention is counted, and what we can't measure.
A complete report on a real business, with its audit findings measured against this scan rather than against a checklist.
Start with your own robots.txt and structured data.
The pre-flight check is free and shows you those findings before you decide anything. The full report is $29.90, once.
Check my siteWe'd like to use Google Analytics to count how many people read this page and how many go on to buy a report — nothing beyond that. It writes a cookie to your browser and sends your IP address to Google, so we have to ask first. Saying no changes nothing about how this site works for you. What it collects.