AI search visibility depends on a sequence of events: a system accesses your content, retrieves it for a question, understands it and decides whether to use it. A useful audit identifies where that sequence is breaking down.
For your business, the practical question is whether search systems can find accurate information about your services and use it to answer a prospective customer. Technical SEO, clear page relationships and credible evidence all contribute.
Different systems crawl and retrieve information in different ways. An answer may combine several sources, and a citation may lead to a visit or simply help someone recognize your business. Each stage needs its own checks.
The Mithril model separates those stages:
- 1Access
Breaks when a crawler is blocked.
The system needs a route to the information before anything else can happen.
- 2Retrieval
Breaks when the page is a weak candidate.
The page needs to match the information the system is looking for.
- 3Understanding
Breaks when the entity is ambiguous.
Names, entities and relationships need to be clear and consistent.
- 4Evidence
Weakens when claims lack useful evidence.
Original experience and verifiable facts give the content a reason to be used.
- 5Selection
Breaks when a stronger source is chosen.
The system compares candidate sources for the particular question.
- 6Citation
A visible reference helps readers identify the source.
Measure visible citations and referral visits separately.
- 7Outcome
Breaks when visibility never becomes a lead.
Connect visibility with qualified inquiries and business results.
Problems at an earlier stage can limit what happens next. Use this guide to identify the weak point, then follow the linked diagnostics to investigate it.
What does it mean to “rank” in AI search?
People understandably use phrases like “rank in ChatGPT” or “rank in AI search.” We use them too. Humans need shorthand. But AI answers do not behave like a universal position 1 through 10 search result.
A system may retrieve several sources for one question. It may generate related searches internally. It may use one page as background and visibly cite another. It may retrieve a page without displaying the source. The answer can also change when the query changes slightly, when fresh information becomes available, or when the underlying model or retrieval system changes.
Bing makes this distinction unusually clear in its AI Performance reporting. Citation counts show that a page was visibly referenced. Bing explicitly says those counts do not represent ranking, authority, importance, or a page’s role inside an answer.
That is why we prefer to think in stages. If you want more AI visibility, identify which stage is weak instead of asking where your imaginary ChatGPT rank went.
Stage 1: Access
Before an AI system can do anything useful with a page, some component of the ecosystem usually needs access to the information. It becomes less obvious once you meet the modern web.
A site can have:
- robots.txt directives
- CDN or firewall rules
- bot-management policies
- JavaScript-rendered content
- content hidden behind interactions
- tracking scripts that replace phone numbers
- third-party widgets
- different mobile and desktop experiences
- canonical tags pointing elsewhere
- noindex directives
- authentication
- rate limits
- geo restrictions
- and approximately seventeen WordPress plugins all convinced they are helping.
Traditional search crawlers already deal with this complexity. AI introduces additional crawler types. OpenAI, for example, separates OAI-SearchBot, which is associated with search discovery and citation, from GPTBot, which site owners can control separately for potential training use.
Cloudflare increasingly separates AI crawler activity into categories such as Search, Training, and Agent. That distinction matters.
“AI bot” is not one job anymore.
Neither should your crawler policy be one big red button. This is why AI Crawlability gets its own guide in this series.
The first rule of machine access
Make essential business facts available through a reliable delivery path. Google can render JavaScript. Google documents a crawl, render, and indexing process for JavaScript pages and uses an evergreen Chromium renderer.
That does not mean every machine visiting the web behaves exactly like Googlebot. Google itself notes that server-side rendering or pre-rendering can still be useful because not all bots can execute JavaScript.
That should change how we think about older web implementations. Consider a law firm website where the actual office phone number is not in the original HTML.
Instead:
- the server returns a placeholder,
- a call-tracking script loads,
- the script detects the visitor source,
- then JavaScript replaces the number.
That may be perfectly appropriate for attribution. But if the canonical business phone number is important for machine understanding, it should exist somewhere reliable too. The same logic applies to service areas, office names, attorney information, FAQs, links, pricing, product details, reviews, and other critical facts.
Observe it.
Learn from it.
Try it on your site
The raw HTML check
This is a diagnostic you can perform without owning a lab coat.
Take a page containing an important business fact. Maybe it is an office address.
Now compare three versions of the page:
- The HTML returned by the server.
- The final rendered DOM in a browser.
- What appears only after a user interaction.
If the address exists in all three, lovely. If it exists only after an API call fires after load, make a note. If it appears only after somebody clicks a tab called “Locations,” make a bigger note.
The point is not that every AI crawler definitely misses the information. The point is that you have created an unnecessary dependency between critical information and a specific rendering behavior.
Our preferred implementation keeps the dependencies clear:
Critical facts should be available as ordinary crawlable content whenever practical.
Boring HTML has survived a remarkable number of technological revolutions.
Stage 2: Retrieval
Being crawlable does not mean being selected for an answer. This is where AI search starts to become more interesting. Google publicly describes its generative Search features as using retrieval-augmented generation and query fan-out.
In plain English, the system can break a complex question into related information needs, retrieve relevant pages from its Search systems, and use those sources to help construct an answer.
Bing exposes something similarly useful through “grounding queries” in AI Performance. These are grouped phrases associated with retrieval activity where a site’s content was cited. This gives us a better way to think about content.
A page does not need to contain the exact prompt somebody typed. It needs to contain useful information for the underlying information need.
That distinction should rescue us from one of the worst possible responses to query fan-out:
creating 100 near-identical pages for every possible wording. Google now explicitly warns against doing that primarily to manipulate generative search. The internet has enough pages called “Best [Service] for [Location] Near Me in 2027.”
Instead, create strong canonical resources and connect them through a useful site architecture.
Stage 3: Understanding
Now the system has a page.
What is it actually about?
This is where keywords become insufficient.
Suppose a page mentions:
- Mithril
- Phoenix
- SEO
- Lumen
- AI search
- technical audits
Those words exist. But machines also need relationships. Mithril is an organization. Mithril provides SEO services. Lumen is a website intelligence project associated with Mithril. Phoenix may describe a market or office location.
A particular person may work for the organization. A service may be available nationally while another is local. This is the entity layer.
You communicate those relationships through many ordinary things:
- clear writing
- consistent naming
- strong About and service pages
- person pages
- location pages
- internal links
- structured data where appropriate
- business profiles
- external references
- and consistency across the places machines encounter the business.
Schema can help clarify this information. Check that the markup represents the page accurately. Google explicitly says there is no special structured data required for its generative Search features. Use structured data because it accurately describes real things on the page and supports the wider search ecosystem.
Choose schema types that accurately describe the page.
Stage 4: Evidence
This may be the most underdeveloped part of most AI search strategies. The internet does not need another page summarizing the same ten facts everybody else summarized. Generative systems are especially good at summarizing commodity information.
Which creates a slightly uncomfortable question for marketers:
If an AI can already generate your article from common knowledge, why does the ecosystem need your article?
Google’s current guidance for generative Search makes this point unusually directly. It recommends unique, non-commodity content and firsthand perspective rather than recycling material that already exists everywhere. That is exactly where we think content strategy is heading.
Useful source material includes:
- original data
- real examples
- case studies
- firsthand observations
- documented workflows
- screenshots
- calculators
- benchmarks
- technical diagnostics
- primary interviews
- expert interpretation
- unusual edge cases
- clear comparisons
- and opinions backed by enough reasoning that somebody can disagree intelligently.
You do not need to become a research university. You do need to contribute something.
A closer look.
Follow the evidence
Myth: “AI content needs to be written in little chunks”
What the evidence says
Choose paragraph length and structure for the reader. Google specifies no universal chunking requirement.
Use headings to organize the argument, tables for comparisons and direct answers where they help. Keep related ideas together and give complex explanations the space they need.
Stage 5: Selection
This is the part marketers most want to reverse engineer.
Why this source?
Why not mine?
There is no public master formula. And anyone offering one should come with a complimentary bag of salt. Selection can depend on the system, query, freshness, source availability, relevance, quality systems, user context, and the information being retrieved.
We can still improve our odds without pretending to know proprietary ranking weights.
A technically available page with a clear purpose, useful information, consistent entities, strong evidence, current facts, and sensible site relationships is simply a better candidate than an ambiguous page full of recycled fluff.
This sounds suspiciously like good SEO. That is because much of it is. Google explicitly says its generative Search experiences remain rooted in its core Search systems. The new work is understanding where AI retrieval, citations, crawler controls, and agent behavior create additional technical considerations.
Stage 6: Citation
Citation is useful. Citation is not the whole game. A page might contribute to an AI answer without receiving the prominent clickable reference you hoped for. A system may synthesize information from multiple places.
Different products expose sources differently. And citations do not automatically equal traffic. Bing is explicit about this too. Its AI Performance reporting distinguishes citation activity from clicks and traditional ranking metrics.
That means “number of citations” should never become the new Domain Authority. Instead, measure citations as one observable signal inside a larger system.
Stage 7: Outcome
This is the part where marketing resumes being marketing.
Did anybody useful discover you?
Did they visit?
Did branded search increase?
Did they convert?
Did a lead say they found you through ChatGPT?
Did the right service page begin receiving traffic from AI referrals?
Did important pages gain citations for topics associated with real revenue?
Visibility without business context is just a very elaborate hobby.
The measurement section of this series gets into:
- Google’s generative AI performance reporting
- Bing AI Performance
- grounding queries
- cited pages
- AI referral traffic
- server logs
- crawler behavior
- synthetic prompt tracking
- conversion tracking
- and the giant caveat attached to every third-party “AI rank tracker”
The short version is:
- Measure what the platforms expose.
- Measure what arrives on your site.
- Use synthetic tracking as directional research.
Keep the prompt sample and collection method visible when reporting the result.
The AI search stack
A practical AI search strategy has several layers.
Crawling and access
- Who can fetch the content?
- What do robots.txt and bot controls permit?
- Are CDN or WAF rules interfering?
- Does critical content require rendering or interaction?
Technical clarity
- Are canonicals correct?
- Are pages indexable where appropriate?
- Are duplicate URLs competing?
- Do internal links expose the structure?
- Are important facts present in accessible content?
Entity clarity
- Who is the organization?
- What services does it provide?
- Who are the people?
- Where are the locations?
- How do those things relate?
Content value
- Does the page contribute anything that did not already exist?
- Does it answer useful questions?
- Does it contain evidence?
- Is it current?
- Would a human actually want to read it?
Source architecture
- Which URL is the canonical home for a fact?
- How do related resources connect?
- Are there clear parent and supporting pages?
Measurement
- Can we observe crawling?
- Can we observe indexing?
- Can we observe citations or grounding queries?
- Can we observe referrals?
- Can we connect any of it to outcomes?
Where does llms.txt fit?
An llms.txt file can guide compatible systems to useful resources, especially on documentation-heavy sites. Its value depends on the consumer and the task.
Google Search ignores the file for visibility and ranking. Prioritize clear pages, accurate facts and reliable access, then evaluate a maintained file for a specific use case. The llms.txt guide explains that decision in detail.
What about AI training?
Choose training permissions separately from search-access goals. OpenAI provides distinct controls for GPTBot and OAI-SearchBot, while Cloudflare distinguishes search, training and agent traffic.
For service businesses publishing freely available educational material, we consider allowing training access where there is no specific reason to restrict it. Publishers and businesses with licensed or proprietary content may make a different choice.
This is a business-policy judgment. Evaluate the rights, content and commercial purpose involved, and document the reason for your choice. Our training-policy guide walks through the questions.
Observe it.
Learn from it.
Try it on your site
The machine-readable business test
Pick one important service and one important location.
Now ask:
- Can I identify the canonical service page?
- Can I identify the canonical location page?
- Does the service page clearly explain who provides the service?
- Does the location page clearly explain what exists there?
- Are the two connected with crawlable links?
- Does the visible information agree with structured data?
- Does the website agree with major business profiles?
- Is the primary phone number understandable even if call tracking changes the displayed number?
- Is important information in the initial HTML or otherwise reliably rendered?
- Is there one obvious current version of each fact?
- Could a machine reasonably build a coherent description of the business from this information?
If the answer is unclear, record the missing information. You found optimization work.
LumenWhat a tool can check
Where Lumen fits
Lumen’s role is to help turn technical evidence into a practical action plan. Begin with the page response, rendered content, links and structured data, then identify the issue that affects a useful business fact.
Keep website findings alongside the separate evidence from Search Console, Bing Webmaster Tools and request logs. State which source supports each conclusion and which questions still need investigation.
What should you do first?
If you are tempted to begin by creating an llms.txt file, buying an AI rank tracker, or commissioning 400 “fan-out optimized” articles, consider starting somewhere less exciting. Make sure important pages are accessible.
Make sure critical facts are not unnecessarily dependent on JavaScript. Make sure the site has sensible canonical URLs. Make sure internal links expose topic relationships. Make sure business information is consistent.
Use structured data accurately. Publish information worth retrieving. Keep important facts current. Understand which crawlers you allow. Measure citations, referrals, and outcomes where real data exists. Then experiment. That is less glamorous than promising a secret ChatGPT ranking factor.
It is also much more likely to survive the next model release.
A closer look.
Follow the evidence
Myth: “SEO is dead because AI answers the question”
What the evidence says
AI changes how answers are presented, while access, useful content and clear site structure remain important.
Review how your pages are discovered and used, then connect observed visibility with visits and inquiries. The channel changes call for better measurement and page quality.
The bigger picture
For years, technical SEO taught websites how to communicate with search engines.
AI search adds new consumers of that information:
- retrieval systems
- generative answer systems
- training crawlers
- browser agents
- and software acting on behalf of people
They do not all behave the same way. They do not all need the same permissions. And they do not all expose the same measurement. That is why we think the next phase of search optimization is not a bag of GEO tricks.
It is disciplined machine-facing information architecture. Make the website easy to access. Make the important facts easy to find. Make relationships obvious. Contribute information worth using. Remove contradictions. Measure what can actually be measured.
Use structured data to clarify supported facts, and assess citation claims against documented evidence.
Sources and primary references
- Google, Optimizing your website for generative AI features on Google Searchdevelopers.google.com
- Google, JavaScript SEO basicsdevelopers.google.com
- OpenAI, Publishers and Developers FAQhelp.openai.com
- Bing Webmaster Tools, AI Performancebing.com
- Cloudflare, AI Crawl Controldevelopers.cloudflare.com

