A customer asks an assistant who to call. Your name comes back, or it doesn’t.
An industry grew around that sentence in about 18 months. It sells files to publish, markup to add, and retainers priced against a problem it rarely defines. The offers arrived faster than the evidence, and the space between them is wide enough to build a business in.
The evidence in this field is thinner than the confidence around it. A single peer-reviewed study exists on what makes content more likely to be cited by a generative engine. The firmest ground is the AI companies’ own documentation, and that documentation is deliberately unspecific about what works — a 20-year-old policy, not an oversight. Everything below was checked against its primary source, and the date of that check is at the foot of the page.
The Three Pillars of AI Visibility
Being visible to an AI assistant rests on three things. Ranking in search, which means traditional SEO matters as much as it ever has. A cohesive narrative across the web, where the facts about your brand agree with each other and the wider web corroborates what you say about yourself. And information that drives a business decision, sitting on your site where a machine can find and read it. An offer addressing fewer than three of those is describing part of the problem.
1. Ranking In Search
Traditional SEO matters as much as it ever has.
2. A cohesive narrative across the web
The facts about your brand agree with each other, and the wider web corroborates what you say about yourself.
3. Information That AI Can Read
It has to be on your site, where a machine can find and read it.
1. Ranking in search
AI visibility is not a new discipline. It is a new consumer of an old one.
Google’s AI answers are assembled through grounding, which the company defines as “relying on our core Search ranking systems to retrieve relevant, up-to-date web pages from our Search index.” The AI does not go looking for you. It asks Search, and Search answers out of the index. If you are not in that index, or you are buried far enough down it that nothing retrieves you, you are not in the pool the answer gets built from.
Google states the conclusion without hedging: “From Google Search’s perspective, optimizing for generative AI search is optimizing for the search experience, and thus still SEO.”
The floor is concrete. A page has to be indexed and eligible to appear in ordinary search with a snippet before it can appear in an AI answer at all. There is no separate AI track, no side entrance, and no file you can publish that moves you up the queue.
Which makes the strangest thing about this market its treatment of search as solved. Every hour of ordinary, unglamorous SEO work — getting pages indexed, fixing the crawl, earning the links that build authority, writing enough depth to be worth retrieving — now pays into two channels instead of one. It was already the work. It has more riding on it than it did.
The weight sits in one half of SEO, and it is not the half that gets sold. Indexable, reachable, technically sound, substantial enough to be worth pulling — that half. The rank-chasing half matters less here than habit suggests: across analyses counting citations rather than pages, a majority of AI citations come from URLs holding no top organic position. Ahrefs, examining 4 million cited URLs in March 2026, put the share drawn from the first page of organic results at roughly 37 percent. A separate BrightEdge analysis landed lower still. Ranking first doesn’t secure the citation. Not ranking at all forecloses it.
This is the part the AI-optimization market walks past. A vendor selling a file to publish is selling a layer that rests on a foundation nobody inspected. If the page isn’t indexed, the file changes nothing.
2. A cohesive narrative across the web
An assistant builds its description of you from every surface it can reach, and it reaches well past your website. Google states that it compiles a local business profile from four kinds of input: crawled web content from the official site, licensed data from third parties, users who contribute facts and photos, and Google’s own interactions with the place.
So a description of your business exists whether you tend it or not. Your site says one thing. Your Google profile says another. A directory nobody has opened since 2019 says a third. None of them is lying — they were written at different times, by different people, for different reasons.
A person reconciles that without noticing. They see the hours disagree, assume the website is stale, and call to check. A machine has no such generosity. It doesn’t infer what you meant, forgive a contradiction, or fill a gap from context it was never given. It assembles what it can reach, and where the pieces disagree, the description it can state with confidence gets smaller. That is how a business ends up described vaguely while a competitor whose facts lined up gets described precisely.
Agreement is only half of it. The other half is whether anything outside your control backs you up.
A claim that appears only on your own website is a claim with a single source. The same claim, repeated in a listing, a review, a trade profile, a mention in someone else’s writing, is a claim the machine can confirm without taking your word for it. That difference shows up in the measurement: in SE Ranking’s analysis of what ChatGPT cites, the strongest correlates were markers of third-party corroboration — who links to you, how much traffic your domain carries, how trusted it is — rather than anything you can add to your own pages. Google’s guidance moved the same way this month, gaining a section on local business details as inputs to generative answers.
The surfaces that decide how you get described are largely ones you don’t own and can’t edit. You can only work to make them agree.
A brand’s job is to be understood: clearly, consistently, over time. We held that position before AI assistants existed, and nothing about them has changed it. Coherence was the standard. What changed is the reader. Coherence stopped being a quality standard and became infrastructure, because the thing reading you now has no capacity to be charitable.
3. Information that drives a decision
Two conditions, and they fail in different ways. The information has to be there. And the machine has to be able to get at it.
Start with what “there” means, because this is the condition almost nothing on the market addresses. A business can rank well, hold every fact in perfect agreement, and still hand an assistant nothing to work with. Consistency isn’t substance. If every surface agrees that you are a full-service firm serving a range of clients, a machine can repeat that with total confidence, and it will still never be the answer to a question anyone actually asked.
The questions people bring to an assistant are decisions in progress. What does this service include, and who is it wrong for. What happens when it goes badly. What determines the price. Should this be repaired or replaced. Pages written to answer those are the ones with something citable in them, because they contain something worth citing.
The one peer-reviewed study on generative visibility found that adding citations, statistics, and direct quotations to a page raised its performance by 30 to 40 percent on the paper’s main measure. That reads like a trick and isn’t one. What made those pages more citable is what makes them more useful to a person deciding: evidence, specificity, claims that can be checked.
Then the second condition, which fails silently. The assistants that fetch pages live — ChatGPT, Perplexity, Claude — come to your server directly rather than through the index. OpenAI recommends allowing its search crawler in robots.txt and allowing traffic from its published IP ranges. The second is where the surprises live, because a firewall or CDN decides it, and whoever wrote the robots.txt file often has no idea what the CDN is doing. Perplexity now publishes a configuration guide for that exact problem.
Being blocked doesn’t produce silence, either. It produces something worse. OpenAI may still surface your link and page title; Perplexity may still keep “the domain, headline, and a brief factual summary.” The system goes on describing you from whatever it gathered elsewhere, and you have no hand in it.
That half of the problem is configuration rather than writing, and it usually belongs to whoever runs the server. It still has to be true before any of the rest counts.
Facts that live only in markup a reader can’t see are facts an assistant may not have. searchVIU built a test page whose price appeared only in structured data and put five systems in front of it; none could use it. Anything that matters belongs in the visible text of the page, not filed where only machines are meant to look.
“Showing up” is three different outcomes
Being named and being the source are different events, and they run on different things.
An assistant can name your business in its answer. It can cite your page as the source it drew from. And — the ordinary case, the one most offers skip past — it can name you while the link points somewhere else entirely.
That third outcome is structural rather than a failure. Google states that it builds a local business profile from four kinds of input: crawled web content from the official site, licensed data from third parties, users who contribute facts and photos, and Google’s own interactions with the place. The picture an assistant holds of you is assembled from surfaces you don’t own, and the link it offers will point at whichever of them it trusts to be checkable.
Access behaves the same way. A site that opts out of OpenAI’s search crawler “will not be shown in ChatGPT search answers, though can still appear as navigational links.” Perplexity says that for a page it has been told not to crawl, it “may still index the domain, headline, and a brief factual summary.” You can be described by a system that has never read you.
Which is why a plan built only around your own website is incomplete before it starts. The website matters. It isn’t where the assembly happens.
The shortcuts, and what each is reaching for
Nearly every shortcut on the market aims at something real. That’s what makes them persuasive, and it’s a better reason to look at them closely than to wave them off.
“Publish a file so the AI can read you”
The popular version is llms.txt — a plain text file at your site root that summarizes your content for language models. It’s a real specification, seriously proposed, and it solves a real problem for a real audience.
That audience is developers. Documentation sites publish these so an engineer working with a coding assistant can hand it a clean version of the docs instead of a sprawl of HTML. Anthropic, Stripe, Cloudflare, Vercel, and Supabase all publish one. Describing that as a convention among publishers is accurate. What isn’t accurate — and this is where the selling happens — is the claim that an assistant consults one of these files when a customer asks about your business. No major provider’s documentation says it does.
Google’s guidance is unambiguous: publishing one “will neither harm nor help your site’s visibility or rankings in Google Search, as Google Search ignores them.” Read that as permission rather than endorsement. Google is saying the file does nothing for them and you may do as you please.
The measurement is harder than the documentation. Ahrefs looked at server logs across 137,000 domains and found that through May 2026, 97 percent of published llms.txt files were never requested at all. Of the requests that did arrive, the largest identifiable share came from SEO auditing tools — software checking whether the file exists, on behalf of people who were told it mattered.
Google can’t quite keep its own story straight here, either. Search says it ignores the file. Chrome added an audit for it to Lighthouse. Read the audit page closely and it never claims a Google system reads one — it’s a checker, not a consumer.
So publish it if you like. It costs nothing. Generate it from whatever already holds your real facts, so it can’t drift and start contradicting you; a stale file is worse than no file, because it manufactures a fresh disagreement in precisely the place you were trying to build agreement. Subverse publishes one on those terms, and it is not why anyone gets recommended.
The formats arriving behind it hold the same shape. Google Cloud’s Open Knowledge Format is still a draft, and the community that has grown around it puts the position plainly in its own FAQ: it’s “designed for your AI systems, not for Google’s crawler.” EntityMap’s specification recommends an ordinary footer link as the most reliable way for crawlers to discover it — new machinery deferring to old.
“Add structured data so it understands you”
Schema markup earns its place in classic search. Rich results, star ratings, the enhanced listings in ordinary Google results — that’s the job, and it does it.
For AI citation, the evidence doesn’t carry the claim. Google states there is “no special schema.org structured data that you need to add,” and separately that structured data “enables a feature to be present, it does not guarantee that it will be present.” The most careful study available compared 1,885 pages that added markup against roughly 4,000 matched pages that didn’t and found no reliable citation lift. Its authors noted that the decline they measured was already underway before any markup appeared, and attributed it elsewhere. The markup didn’t damage those pages. It just didn’t do the thing it was sold to do.
There’s one genuine exception and it’s narrow. Google’s Merchant Center accepts feed attributes described as “primarily intended for use in conversational experiences such as AI mode in Google Search.” That is structured data feeding an AI surface — but it’s a product feed submitted to Google, not markup added to a page, and it applies to ecommerce. That distinction settles most of the argument on its own.
The figures circulating in support of schema as an AI lever — a 27 percent visibility gain, a claim that one assistant cites structured content 30 percent more often — trace back to nothing. Follow them and you find blogs citing blogs.
“Get your brand mentioned everywhere”
Corroboration is real, which makes this the most plausible of the three.
Google’s position is specific: “Seeking inauthentic ‘mentions’ across the web isn’t as helpful as it might seem.” Its spam policies name attempting to manipulate generative AI responses in the opening line of what spam is. The objection isn’t to mentions — mentions are among the strongest correlates in that same analysis. The objection is to manufacturing them, and the systems doing the reading are built to discount exactly that.
The pattern underneath
Set the three beside each other and the same move appears in all of them. A structural question — can a machine assemble a confident description of this business — gets answered with an artifact. Publish a file. Adopt a format. Buy some mentions. Each is cheaper than the work it substitutes for, and each is far easier to sell, because it can be delivered, invoiced, and shown.
None of this is fraud. It’s an industry building supply ahead of demonstrated demand. It’s also familiar. Branding spent two decades answering a structural problem with more output — more campaigns, more collateral, more visibility — and produced more of the condition it was hired to resolve.
How you’d know any of it was working
Nobody can promise that an assistant will recommend you. Anyone who does is describing a system they have no access to.
Google put this in writing recently, and the sentence rewards a second reading: “Be wary of third-party tools that promise ranking success or claim to use ‘internal’ Google metrics. No third-party tool has access to our internal ranking or AI systems.” That covers every vendor in this market. It covers us.
What can honestly be done is narrower and duller than what’s advertised. You take the questions a real customer would ask, put them to the assistants your customers actually use, and record what comes back — who gets named, who gets cited, where you appear, and where a competitor appears instead. Then you do it again on a schedule and watch whether the line moves. That isn’t privileged access. It’s observation, repeated, which is the only instrument anyone in this market has.
Google now reports some of this directly in Search Console, which beats an intermediary reselling an estimate of it.
The rest is the work: being findable, being coherent, being worth citing. None of that is new. What’s new is the reader — one with no capacity to infer, forgive, or extend you the benefit of the doubt. Everything that used to be good practice is now load-bearing.
None of it arrives at once. These are three places the same picture can break, not a checklist to clear before anything counts, and knowing which one is broken for you is most of the beginning.
This is the work we do. If that’s useful, it’s described here. If it isn’t, everything above stands without it.
What this is based on
Every claim above was checked against its primary source on 19 July 2026. Dates below are each source’s own — several of these pages moved during the month this was written.
The AI companies, in their own documentation
- Optimizing for Google’s AI features — Google Search Central, updated 2026-07-10. “Still SEO,” the definition of grounding, the
llms.txtguidance, inauthentic mentions, micro-chunking, and the warning about third-party tools. - AI features and your website — updated 2025-12-10. The eligibility floor, and “no special schema.org structured data that you need to add.”
- Search generative AI control — the Search Console setting, and confirmation that inclusion is the default.
- Structured data general policies — updated 2026-07-10. Structured data “enables a feature to be present, it does not guarantee that it will be present.”
- Google Search spam policies — updated 2026-05-15. Manipulating generative AI responses named in the opening definition.
- How Business Profile information is sourced — the four kinds of input.
- Merchant Center conversational attributes — the feed attributes intended for AI surfaces.
- Lighthouse
llms.txtaudit — Chrome’s audit, which checks for the file without claiming anything reads it. - OpenAI: overview of crawlers — the opt-out consequence, and the recommendation to allow both the crawler and the published IP ranges.
- OpenAI: publishers and developers FAQ — what can still surface from a blocked page.
- Perplexity crawler documentation — bot behaviour, and the firewall configuration guidance.
- How Perplexity follows robots.txt — updated 2026-07-16. What remains indexed when a page is blocked.
- Anthropic: web search tool — when Claude searches, and that citations are always on.
Peer-reviewed
- GEO: Generative Engine Optimization — Aggarwal et al., KDD ’24. A controlled experiment across 10,000 queries. The only peer-reviewed study cited here.
Independent analysis
Design and sample are given because they matter. None of these is a controlled experiment; all observe patterns rather than test causes.
- 97% of
llms.txtfiles never get read — Ahrefs, published 2026-06-15 using May 2026 data. Server-log analysis across 137,210 domains. The authors note that “fetched” and “acted on” are different questions, and that their adoption figure is an upper bound. - AI Overview citations and the top 10 — Ahrefs, 2026-03-02. 863,000 keyword SERPs, 4 million cited URLs. Observational.
- Schema markup and AI citations — Ahrefs, 2026-05-11. 1,885 pages that added markup against roughly 4,000 matched controls. The strongest design of the non-academic work cited here, and its authors attribute the decline they measured to broader trends rather than to schema.
- Rank overlap after 16 months of AI Overviews — BrightEdge, 2025-09-18. Sample size not disclosed, which is why it appears here without a figure.
- What correlates with being cited by ChatGPT — SE Ranking, 2025-11-24. 216,524 pages. Measures ChatGPT specifically, not AI Overviews. The collection window is not disclosed.
- What ChatGPT, Claude, Perplexity and Gemini really see — searchVIU, tested 2025-10-30. A single purpose-built test page. An existence proof, not a rate.
The specifications themselves
- llms.txt — the specification, stable since September 2024.
- Open Knowledge Format FAQ — maintained by the community around the format, not by Google.
- EntityMap specification — including its own note that a footer link remains the most reliable route for crawlers.
A note on the evidence
One peer-reviewed study exists in this area. Most of what circulates is vendor research — some of it careful, much of it undated, and nearly all of it produced by companies selling something adjacent to the finding. That includes several sources above, and it is why each carries its date, sample, and design.
The AI companies’ own documentation is the firmest ground available, and it is deliberately non-specific about what works. That is a twenty-year-old policy rather than an oversight: a published playbook gets gamed. Reading meaning into what Google declines to say is the same error the vendors make, and it isn’t made here.