See What AI Says
BENCH-2026-08-34 · scanned 28 August 2026

We scanned 34 market leaders. Most of them haven't done it either.

If you run a small business, the honest first question about all of this is whether it's real or whether someone is selling you a new anxiety. So before we measured any customer, we measured the top of the market: 34 well-known companies, checked for exactly the things the GEO industry tells you to go and do. Here is everything we found, including the parts that argue against buying anything.

What we actually did

On 28 August 2026 we requested a fixed list of URLs from 34 websites and read what came back. Zero AI calls, zero cost, no interpretation in the middle — this is a file-and-header scan, the same mechanical checks that run against your own site in the audit section of a report. That is also the limit of it, and we come back to that below.

The 34, in three groups:

For each one we asked for 10 AI-facing files (llms.txt, agents.md, ai.txt, an agentic-discovery sitemap, .well-known/mcp.json and others), read robots.txt and checked it against 22 named AI crawlers, and pulled the homepage to count its structured data, its h1 elements and how many words of readable text the HTML contains before any JavaScript runs.

Three of the 34 refused us. Warby Parker returned 403, Patagonia 404 and away.com dropped the connection — bot mitigation reacting to a datacentre IP address, almost certainly, not a statement about their sites. Their homepages are therefore unmeasured, not zero, and every count below is out of the 31 homepages and 30 robots.txt files we could actually read. Roborock's robots.txt also failed to load.

Finding 1: the file you are being told to write, most of them don't have

llms.txt is the flagship deliverable of a lot of GEO packages. Across the 34:

FileFound on
/llms.txt22 of 34
/llms-full.txt12 of 34
/agents.md (or /AGENTS.md)9 of 34
/sitemap_agentic_discovery.xml7 of 34
/.well-known/mcp.json4 of 34
/.well-known/ai-plugin.json1 of 34
/ai.txt0 of 34
/.well-known/agents.md0 of 34
/mcp0 of 34

It splits sharply by industry. 13 of the 16 software companies serve an llms.txt; 4 of the 8 DTC brands and 5 of the 10 China-based brands do. Companies whose customers are developers publish developer files. That is a less exciting explanation than "the leaders are ahead of you", and it fits the data better.

Two of these deserve saying out loud because they get sold anyway: nothing in this sample served an ai.txt, and Google's own documentation says it does not read llms.txt. A file is not a switch.

Finding 2: over half of the "AI introductions" were written by the shopping platform

Nine sites carry an agents.md — the file that tells an AI agent who you are and how to deal with you. We read all nine. Six of them are the same file.

Allbirds, Glossier, Gymshark, Brooklinen, EcoFlow and UGREEN each serve a 4,129–4,311 byte document headed # Agent Instructions — {brand}, and each of the six points an AI agent at shop.app — Shopify's own checkout — exactly eight times. Compared character by character, Allbirds' copy and UGREEN's differ only in the brand name and one policy link. Five of the six also serve an llms.txt containing exactly five links.

We have seen this file before. It was also on the site of a one-person handmade skincare business we scanned while building the method — same template, same eight pointers to someone else's checkout. So a large athleisure brand and a soap maker working out of a spare room currently introduce themselves to AI with the same paragraph, and neither of them wrote it.

Three sites wrote their own. Shopify's runs to 12,522 bytes, Zapier's to 1,205, and Anker's to 6,610 — Anker being the only brand in the sample that took the platform default and rewrote it, dropping shop.app from it entirely and adding sections on regional markets and canonical product links.

The point for you is not that these files are worthless. It's that having one and having written one are different states, and any audit that reports the first as a green tick is measuring your hosting provider.

Finding 3: the hand-written ones are mostly a table of contents

Where a company clearly did write its own llms.txt, the dominant format is a plain annotated list of its own pages. Stripe's carries 288 links, HubSpot's 201, SHEIN's 190. The platform-generated ones carry five.

That is a description of what they did, not evidence that it worked. We did not measure whether any of these files changed how an assistant talks about the company, and we are not aware of anyone who has published that measurement either.

Finding 4: almost nobody is blocking AI crawlers

This is the question small-business owners ask us most: should I block the AI bots? Of the 30 robots.txt files we could read, only 7 mention any of the 22 AI crawlers we checked, and only 3 block one:

SiteAI crawlers fully disallowed
Canva12
Figma7
Notion2
The other 27 readable sitesnone

The three who do block are all companies whose product is a design or document canvas, i.e. companies with an obvious training-data reason of their own. Everyone else lets them in. If someone has told you to shut the AI crawlers out as a defensive measure, the top of the market is doing the opposite.

Finding 5: the basics aren't universal at the top either

This is the part that changed how we write reports. Out of the 31 homepages we could read:

CheckHead-of-market sites passing
At least one h1 on the homepage21 of 31
Any JSON-LD structured data at all20 of 31
Organization schema14 of 31
AggregateRating schema2 of 31
Review schema0 of 31

Ten of these companies have no h1 on their homepage. Eleven publish no structured data whatsoever. And review markup — a standing item on essentially every SEO checklist ever written — appears on two sites out of 31 in the form of an aggregate rating, and on none at all as individual reviews.

Readable text is uneven too. The median homepage served 1,059 words of text in its HTML before any JavaScript ran; the largest served 40,745. Four served under 300: UGREEN 72 words, Gymshark and Narwal one word each, Temu none. Their pages plainly work for humans in a browser — but an AI system reading the raw HTML sees close to nothing, and we did not run JavaScript, so all this measures is what arrives in the first response.

So how far off is a small business?

Nearer than the marketing suggests, on the mechanical checks at least. The column on the right is the five small businesses we ran the full method against while validating it. We don't name them and we don't publish any individual result — that rule applies to everyone who buys a report, which is why the sample report is pseudonymised too. Five is a tiny sample and the numbers below should be read as a sketch, not a rate.

CheckHead of marketFive small businesses
h1 on the homepage21 of 311 of 5
Organization schema14 of 312 of 5
Review / AggregateRating schema2 of 310 of 5
robots.txt naming any AI crawler7 of 300 of 5

Read the rows in order. The first is a genuine gap and worth ten minutes of somebody's time. The second is a gap against a majority that isn't itself a majority. The third and fourth are not gaps at all — nobody is doing them, at any size, and a report that listed them as your failings would be padding.

That is now a rule in how we write: a check has to be compared against this benchmark before it can be called a problem. Below the benchmark is a problem. Level with it is at most an opportunity. Above it gets said out loud as a win. It removed two items from our audit list outright, and it is the reason your report will not be a 60-line sheet of things you have not done.

What this study does not show

It is a file scan. It measures what these companies have deployed, and that is all it measures.

What to take from this if you run a small business

Three things, and none of them is "buy something".

  1. You are not as far behind as the pitch implies. Over half of these companies have no dedicated AI file beyond an llms.txt, two thirds of the ones that do had it written by their shopping platform, and eleven of them publish no structured data at all. Anyone quoting you thousands to catch up with the leaders should be asked which leaders, and shown this.
  2. Don't block the AI crawlers. 27 of the 30 sites we could read let all of them through. This is the single cheapest thing on the page: go and look at your own robots.txt, for free, right now.
  3. Files are not the interesting question. Everything on this page is infrastructure, and infrastructure is the part you can copy in an afternoon. What it cannot tell you is whether an assistant says your name when a customer asks — and on that, being at the top of the market clearly does not depend on having any of these files, because several of these companies have none.

The only way to know the second thing is to ask the assistants, in your category, and look at what comes back. That's the measurement, and you can start it without paying us.

How the real measurement works →

23 questions across 3 AI surfaces, how a mention is counted, and what we can't measure.

See a benchmark comparison in place →

A complete report on a real business, with its audit findings measured against this scan rather than against a checklist.

Start with your own robots.txt and structured data.

The pre-flight check is free and shows you those findings before you decide anything. The full report is $29.90, once.

Check my site

We'd like to use Google Analytics to count how many people read this page and how many go on to buy a report — nothing beyond that. It writes a cookie to your browser and sends your IP address to Google, so we have to ask first. Saying no changes nothing about how this site works for you. What it collects.