The llms.txt Myth: What Actually Gets Brands Cited by ChatGPT
The llms.txt Myth
What actually gets brands cited by ChatGPT — a 6-brand teardown
Every AEO checklist published this year says the same thing in the same order: write an llms.txt file, structure your content for AI crawlers, and the citations will follow. We wanted to know if that’s actually true. So we picked six real e-commerce brands, asked ChatGPT the questions real customers ask, and checked — line by line — what got cited, what got ignored, and why.
The honest answer surprised us. llms.txt turned out to be one of the least important variables in the entire experiment. What mattered more was messier, more human, and much harder to fake.
The Myth We Set Out to Test
The pitch behind most AEO service offerings right now goes roughly like this: “Google can’t read your JavaScript-rendered site, and neither can ChatGPT. Add an llms.txt file, structure your content in clean markdown, and AI engines will start citing you.” It’s not a crazy idea — it’s just incomplete, and treated as gospel it can lead brands to spend budget on the wrong thing.
We tested it directly. Six brands, three prompt types each, across ChatGPT, with cross-checks in Google’s own AI Overview:
- Warby Parker (eyewear) — has an llms.txt
- Ace & Tate (eyewear competitor, surfaced organically in testing)
- Glossier (beauty) — has an llms.txt
- Chewy (pet retail) — no llms.txt
- Tiffany & Co. (luxury jewelry) — no llms.txt
- Bombas (apparel) — no llms.txt
The prompts, repeated for each brand:
- A direct factual question: “What is [Brand]’s return policy?”
- A comparison question: “What are the best [category] brands like [Brand]?”
Myth #1: “No llms.txt Means No Citation”
This is the easiest myth to kill, because the data is unambiguous. Three of our six brands — Chewy, Tiffany & Co., and Bombas — have no llms.txt file at all. We confirmed the 404 on all three directly.
What we found
| Brand | llms.txt? | Return-policy citation |
| Chewy | No (404) | Clean, single-source, own domain |
| Tiffany & Co. | No (404) | Clean, single-source, own domain |
| Bombas | No (404) | Clean, single-source, own domain |
| Glossier | Yes | Clean, single-source, own domain |
| Warby Parker | Yes | Partial — shared with 3rd-party sites |
[VERIFIED] Confirmed via direct 404 checks at each domain’s /llms.txt path and via visible utm_source=chatgpt.com parameters on the cited URLs, screenshotted during live testing.


Look closely at that last row. Warby Parker is the one brand in this table that got a worse result than three brands with zero AI-crawler infrastructure. It has a well-built llms.txt file — dedicated markdown pages for prescriptions, company history, manufacturing — and it still got its return-policy citation diluted by third-party aggregator sites instead of a clean, single-source answer.
“Having an llms.txt file isn’t the same as having the right llms.txt file.”
waqar ahmed
The reason is simple once you see it: Warby Parker’s llms.txt doesn’t include a page about returns or refunds at all. Glossier’s does — with dedicated markdown files for “What’s your return policy?”, “How do I place a return or exchange?”, even edge cases like gifted items and TikTok Shop returns. When the exact question a customer asks maps to a page that explicitly answers it, the citation is clean. When it doesn’t, ChatGPT goes looking elsewhere — and “elsewhere” is often a third-party aggregator that happens to have crawlable, well-structured content on the topic you left uncovered.
Meanwhile Chewy, Tiffany, and Bombas — none of which have any llms.txt — got clean, single-source, own-domain citations for the same type of question. Why? Because each of them has an unambiguous, crawlable page that directly answers “what is your return policy,” sitting on their main domain in ordinary HTML.
The corrected rule
llms.txt is not the deciding factor for basic factual citation. What matters is whether a clear, crawlable page answering the exact question a customer would ask exists anywhere on the domain. A structured AI-crawler file is one way to signal that — but it only helps if it actually covers the query, and it isn’t required if your ordinary site architecture already does the job.


What Google Itself Says About This
This isn’t just our own finding. In a June 2026 Search Off The Record episode, Google’s John Mueller addressed llms.txt directly — and confirmed the mechanism behind what we saw in testing.
Mueller revealed that llms.txt was never designed for discovery in the first place:
“The idea was really not to create something that makes it easier for search engines or LLM systems to discover all of your content… I think the aspect of using this as a way to optimize for Discovery by AI systems or Discovery by search systems, that doesn’t make any sense at all.”
He also explained why brand-authored llms.txt files carry limited weight with AI systems to begin with:
“You’re basically telling these systems, like, I have the best website ever… So in an LLM system, it — by design — can’t trust what is here as a way of differentiating between different websites.”
That lines up with what we found: discovery and citation still run through ordinary, crawlable HTML — not through a self-declared llms.txt file. Chewy, Tiffany & Co., and Bombas all got clean citations without one; Warby Parker had a well-built one and still got diluted, because the file wasn’t the thing doing the work.
[VERIFIED] Quotes from John Mueller, Google Search Relations, via Search Engine Journal, “Google Exposes The Fundamental Flaw Of LLMs.txt,” June 23, 2026 (Roger Montti).
Source: searchenginejournal.com/google-exposes-llms-txt-flaw/579814/
Myth #2: “Getting Mentioned Is the Goal”
Most AEO advice stops at “get your brand mentioned in AI answers.” Our testing surfaced something more specific and more useful: there’s a tier above being mentioned. In every comparison-style query we ran, ChatGPT didn’t just list competitors — it elevated one or more of them with images, narrative framing, and deeper detail. Being in the list and being in that elevated tier are two very different outcomes, and they’re won differently.
We expected one universal mechanism behind that elevation. Instead we found four distinct shapes, plus a fifth that showed up when we extended the test — five different “games” a brand can be playing depending on its category.
The Five Citation Shapes
Here is the actual taxonomy that emerged from six brands and two categories of query. This is the part worth screenshotting.
Shape 1 — The Single Winner (Eyewear, Beauty)
In Warby Parker’s comparison answer, ChatGPT listed nine competitors in a clean table — brand names hyperlinked, exact dollar price ranges. But only Ace & Tate got the deep-dive treatment below the table: three images, a narrative paragraph, and a citation to “industry rankings.” Ace & Tate has no llms.txt at all.
We repeated the same test for Glossier and got the identical shape: a comparison table of nine-plus brands, then a single deep-dive on Rhode — three images, a founder callout on Hailey Bieber, a narrative paragraph.
What actually predicts the winner
We initially assumed the winner would be whichever competitor had the most third-party “best of” coverage. That held for eyewear but broke down for beauty:
- Ace & Tate (eyewear): dense, fresh (2026-dated), unambiguously positive roundup coverage — ranked #2 in a June 2026 “best glasses” listicle, multiple Reddit threads validating quality, press coverage positioning it alongside Chanel and Dior.
- Warby Parker (eyewear, did NOT get the deep-dive slot): also has substantial roundup coverage — but its Reddit thread was framed as skeptical debate (“is it actually better than Luxottica-made frames?”), not pure endorsement.
- Rhode (beauty): got the deep-dive slot despite category-specific roundup coverage that was thinner and more inconsistent than Ace & Tate’s — several fresh 2026 “brands like Glossier” listicles exclude Rhode entirely, and Google’s own AI Overview for the same query doesn’t name it.
[VERIFIED] Roundup dates, rankings, and Reddit thread framing confirmed via direct Google searches during testing, screenshotted live.


The honest conclusion: recency and framing of off-site mentions predicted the winner cleanly in eyewear, where the competing signal was other roundups. In beauty, it didn’t — Rhode’s advantage looks more like general fame and celebrity association (Hailey Bieber) overriding a thinner, more contested category-specific consensus. Two categories, two different reasons for the same citation shape. That’s a more useful finding than a single tidy rule, because it tells you the mechanism you’re up against depends on what your category actually runs on.
Shape 2 — The Decision Matrix (Pet Retail)
Chewy’s comparison answer broke the single-winner pattern entirely. Every major competitor — Petco, PetSmart, Amazon, Only Natural Pet, Hollywood Feed, Tractor Supply Co. — got full profile treatment: images, narrative, “best for” bullets. The answer closed with an explicit decision matrix: “Most similar to Chewy: Petco. Best in-store experience: PetSmart. Fastest shipping: Amazon. Best natural products: Only Natural Pet.” No single brand was crowned.
When we checked the off-site landscape for pet retail, the reason became clear: there is no dense “best pet retailer” roundup ecosystem the way there is for eyewear or beauty. Google’s own AI Overview even named a different top-three (Petco, PetFlow, Walmart) than ChatGPT did. The Reddit thread we found was people asking for alternatives, not endorsing a winner. The listicles that did exist were actually dog-food product roundups, a different content category entirely.
“Where there’s no off-site consensus to draw from, ChatGPT doesn’t invent one — it defaults to an even-handed comparison guide instead.”
Waqar ahmed


Shape 3 — Segmented by Use Case, Sourced from Wikipedia (Luxury Jewelry)
Tiffany & Co.’s comparison answer took a third shape entirely: recommendations segmented by use case — “For engagement rings: Harry Winston, Cartier, Graff,” “For everyday luxury: David Yurman, Bulgari,” “For heirloom pieces: Van Cleef & Arpels, Mikimoto.” No single winner, no full decision matrix — a curated set of answers depending on what the customer actually wants.
The most interesting detail here wasn’t the format — it was the sourcing. Clicking a competitor’s name didn’t open an external site. It opened an in-app ChatGPT panel with images and a short bio sourced explicitly from Wikipedia, not the brand’s own marketing copy, not a press feature, not Reddit.
When we checked the off-site landscape, luxury jewelry turned out to have the strongest, cleanest consensus of any category we tested. Google’s AI Overview named the exact same trio ChatGPT did — Cartier, Van Cleef & Arpels, Harry Winston — and multiple fresh 2026 roundups reinforced it almost word for word. The likely explanation: heritage jewelry houses have decades of stable, largely undisputed prestige rankings, unlike younger DTC brands where category consensus is still forming or actively contested.

Shape 4 — Community Consensus, No Imagery (Apparel/Basics)
Bombas’s comparison answer was the plainest of the five. A first table of sock-specific competitors (Darn Tough, Smartwool, Stance) with no pricing column. Then a callout explicitly citing Reddit: “community recommendations consistently favor Darn Tough, Feetures, and American Trench.” Then a second, purely text list for apparel beyond socks. Then a compact micro-matrix (“Best overall quality: Darn Tough. Best athletic: Feetures.”). No brand — not even the highest-ranked competitor — got the image-heavy deep-dive treatment that eyewear, beauty, or jewelry brands received.
This Reddit-citation pattern is worth its own deep dive — we tested it directly in The Reddit Paradox: Why ChatGPT Cites the Platform It Relies on Least, running the same query across ten brands to find out exactly when Reddit gets cited versus just quietly consulted.
Instead, the closing paragraph looped back to Bombas’s own buy-one-donate-one giving model as its defining differentiator, citing an industry publication rather than crowning a competitor. Socks and basics, like pet retail, don’t have a “vibe” roundup culture — but unlike pet retail, there was just enough Reddit-driven durability discussion to produce a specific, named community consensus, without the visual investment seen in more image-driven categories.


What the Five Shapes Have in Common
| Category | Winner mechanism | Deciding signal found off-site |
| Eyewear | Single winner, image deep-dive | Fresh, unambiguous roundup consensus |
| Beauty | Single winner, image deep-dive | Celebrity/press fame overriding thin consensus |
| Pet retail | Full decision matrix, no winner | No roundup culture exists for the category |
| Luxury jewelry | Segmented by use case, Wikipedia-sourced | Strongest, cleanest roundup consensus found |
| Apparel/basics | Text-only community consensus | Reddit-specific durability discussion, no imagery |
The pattern underneath the pattern: ChatGPT isn’t applying one “best brand” algorithm uniformly. It’s reading whatever off-site signal actually exists for that category and reproducing its shape. Categories with strong prestige-driven “best of” culture (jewelry, and to a degree eyewear) get single-winner or segmented treatment sourced from that culture. Categories with no such culture (pet retail) get an even-handed functional comparison instead. Categories with active but informal community discussion (Reddit-heavy apparel/basics) get a community-consensus citation with no visual investment.
What This Means for Your Brand
If you’re building an AEO strategy off a generic checklist, here’s what this teardown suggests you check instead:
- Audit whether your site has a clear, crawlable, single-purpose page answering the exact questions customers actually ask — before worrying about llms.txt at all.
- If you do build an llms.txt file, treat it as a coverage exercise, not a checkbox. List the actual queries your support team gets, and confirm each has a dedicated page in the file.
- Understand what “winning” looks like in your specific category before investing in off-site PR. If your category runs on prestige roundups (jewelry, eyewear), fresh, unambiguous “best of” placements matter enormously. If it runs on Reddit-style community trust (apparel, basics), forum presence matters more than press. If your category has no roundup culture at all (pet retail, and likely many B2B or utility categories), a single “best” isn’t being decided — being present and accurately represented across competitors may matter more than trying to be crowned.
- Watch for the hyperlink/pricing-specificity signal. Answers with hyperlinked brand names and exact pricing appear to be more likely browsing-backed; plain-text answers with vague pricing symbols may be running on older trained knowledge. If your brand’s pricing or positioning has changed recently, a plain-text answer is a sign the model may be working from stale information.
The Honest Limits of This Teardown
This is six brands, one AI engine as the primary test surface, and a snapshot in time. Google’s AI Overview showed meaningfully different competitor sets than ChatGPT in more than one category — model-specific and even session-specific variation is real, and any brand relying on a single test run is looking at one data point, not a trend.
[DISPUTED] The underlying mechanisms proposed here (recency weighting, sentiment framing, category-culture defaults) are our best explanation for the observed patterns, not confirmed model behavior — treat as informed hypothesis, not documented fact.
The Findings, Recapped
- llms.txt is neither necessary nor sufficient. Chewy, Tiffany & Co., and Bombas all got clean, single-source citations with zero AI-crawler infrastructure. Warby Parker had a well-built llms.txt file and still got its citation diluted — because the file didn’t cover the one query that mattered.
- What actually predicts a clean citation is coverage, not the file format. A brand gets cited cleanly when a page exists anywhere on its domain that directly answers the exact question being asked — in llms.txt, in ordinary HTML, or both.
- Getting mentioned and getting recommended are different outcomes. Every comparison query produced a list of competitors, plus a smaller, elevated tier that got images, narrative framing, and deeper detail. Being in the list doesn’t mean being in that tier.
- There are (at least) five distinct shapes that elevated tier can take. Single Winner (eyewear, beauty), Decision Matrix (pet retail), Segmented by Use Case with Wikipedia sourcing (luxury jewelry), and Community Consensus with no imagery (apparel/basics) — each tied to how that category is actually discussed off-site.
- What wins the elevated slot is category-specific, not universal. Fresh, unambiguous roundup consensus won it in eyewear and (most cleanly) in luxury jewelry. Celebrity and press fame won it in beauty despite thinner category consensus. Reddit-specific durability discussion won it in apparel. In pet retail, no such consensus existed at all — so ChatGPT gave every competitor equal treatment instead of picking a winner.
- Hyperlinks and price specificity are a tell. Answers with hyperlinked brand names and exact dollar figures behaved as if they were live-browsed and search-grounded. Answers with plain text and vague pricing symbols behaved as if they were pulled from older trained knowledge — a useful signal for judging how current any given AI answer actually is.
- The right AEO move depends on knowing which game your category is playing. Prestige-driven categories need fresh, unambiguous “best of” placements. Community-driven categories need genuine Reddit and forum presence. Categories with no roundup culture at all need accurate, complete representation more than a crowned “best.”
Closing
The myth we set out to test — that llms.txt is the master key to AI citation — didn’t survive contact with six real brands and a browser. What we found instead is more useful, if less tidy: getting cited well depends on matching your content and off-site presence to the specific way your category gets discussed, recommended, and ranked. That’s not a five-minute fix. But it’s a real one.

One Comment