Skip to content

Your client portal is here, with your reports, updates and tasks under one roof. Sign in to Lumen

Search + AI

The Technical AI Search Audit: What We Actually Test

A practical technical AI search audit covering crawler access, rendering, robots, canonicals, entities, schema, retrieval readiness, evidence and measurement.

In this article 23 sections

A technical AI-search audit checks whether important business information is accessible, clear, consistent and supported. It produces findings your team can investigate and act on.

Our approach follows seven observable stages:

The Mithril modelFrom access to an outcomeSeven stages, from a crawler reaching the page to a lead on the phone.
  1. 1
    Access

    Breaks when a crawler is blocked.

    The system needs a route to the information before anything else can happen.

  2. 2
    Retrieval

    Breaks when the page is a weak candidate.

    The page needs to match the information the system is looking for.

  3. 3
    Understanding

    Breaks when the entity is ambiguous.

    Names, entities and relationships need to be clear and consistent.

  4. 4
    Evidence

    Weakens when claims lack useful evidence.

    Original experience and verifiable facts give the content a reason to be used.

  5. 5
    Selection

    Breaks when a stronger source is chosen.

    The system compares candidate sources for the particular question.

  6. 6
    Citation

    A visible reference helps readers identify the source.

    Measure visible citations and referral visits separately.

  7. 7
    Outcome

    Breaks when visibility never becomes a lead.

    Connect visibility with qualified inquiries and business results.

TechnicalAccessible, readable pages
ContentClear, useful evidence
OutcomeA useful next action

We inspect the website, its responses and the available reporting data. The result is a prioritized set of issues, each with evidence, an explanation and a recommended next step.

Part 1: Access

First question:

Can relevant machines reach the information?

We review:

  • robots.txt
  • crawler-specific directives
  • meta robots
  • X-Robots-Tag
  • CDN/WAF controls
  • bot-management policies
  • rate limits
  • authentication
  • status codes
  • redirects
  • server responses
  • and known crawler behavior where logs are available

We distinguish:

Do not report:

“AI bots blocked”

without saying which bot, why it matters, and whether the state is intentional.

Mithril LabsTry it.
Observe it.
Learn from it.

Try it on your site

Access contradiction check

Compare:

If robots says Allow and the WAF returns 403, the infrastructure wins. If robots blocks the page, a crawler may never read the noindex on that page. If an AI-search crawler is blocked while visibility in that product is a business goal, document the mismatch.

The finding should connect technical state to intent.

Part 2: Raw HTML versus rendered page

Next:

What does the server actually send?

Compare:

  • initial HTML
  • rendered DOM
  • interaction-dependent content
  • third-party widgets

Important checks:

  • Does the primary service exist in raw or reliably rendered content?
  • Are important links crawlable?
  • Are contact details stable?
  • Does DNI replace phone numbers?
  • Are location details injected?
  • Is schema client-side?
  • Do FAQ or tab contents require interaction?
  • Does navigation exist before scripts initialize?

This is one of the strongest places Lumen can contribute. A browser screenshot cannot answer these questions.

Part 3: Discovery and site architecture

Important pages need paths.

Review:

  • internal links
  • navigation
  • breadcrumbs
  • sitemaps
  • orphan pages
  • crawl depth
  • redirecting internal links
  • parameter traps
  • faceted navigation
  • pagination
  • and broken links

Then map page roles.

Which URL is the pillar?

Which URL owns the service?

Which URL represents the office?

Which URL represents the person?

Which supporting pages deepen the topic?

A website can have thousands of indexable pages and still have no coherent architecture.

Part 4: Canonical truth

Find competing versions.

Review:

  • rel=canonical
  • redirects
  • sitemap URLs
  • HTTP/HTTPS
  • www/non-www
  • parameters
  • old slugs
  • campaign copies
  • staging copies
  • duplicate content
  • near-duplicate content
  • and conflicting internal links

Then review facts.

Does one obvious source own:

  • business name
  • phone
  • address
  • hours
  • people
  • services
  • pricing
  • policies
  • credentials

AI search makes stale canonical truth more visible because generative answers can surface the wrong version confidently. Give systems fewer wrong versions.

Part 5: Entity clarity

Inventory the things that matter.

  • Organization
  • People
  • Locations
  • Services
  • Products/tools
  • Credentials
  • Key content assets

For each entity, ask:

  • Is there a canonical page?
  • Is the name consistent?
  • Are relationships explicit?
  • Does structured data reflect the page?
  • Do internal links connect related entities?
  • Do major external profiles agree?

This is especially important for local and professional-service businesses.

Part 6: Structured data

We do not score schema quantity.

We evaluate:

  • validity
  • appropriateness
  • visible-content alignment
  • entity relationships
  • canonical identifiers
  • duplicate/conflicting nodes
  • JavaScript dependencies
  • location accuracy
  • person accuracy
  • and eligibility for relevant search features

Remember:

Google says there is no special schema required for generative Search.

Schema is an encoding layer. Use it accurately.

Part 7: Retrieval readiness

Now ask whether important pages are useful retrieval candidates.

  • Does each page have a clear topic?
  • Does it answer the main information need?
  • Are important facts explicit?
  • Can the page stand alone?
  • Does it contain relevant evidence?
  • Is critical information surfaced reasonably early?
  • Is the content current?
  • Does the URL have a unique reason to exist?
  • Are there stronger duplicate pages?

Bing’s current guidance gives us useful observable principles here:

  • focused URLs,
  • explicit facts,
  • clear headings,
  • current information,
  • crawlable links,
  • and independent verifiability.

These do not guarantee a citation. They create better source candidates.

Mithril LabsTry it.
Observe it.
Learn from it.

Try it on your site

The “why this page?” test

For each priority URL, finish:

“This page should be retrieved because...”

Weak

It targets the keyword.

Better

It is the canonical page explaining Mithril’s technical AI-search audit process, includes the full methodology, links to supporting diagnostic guides, and contains information not repeated elsewhere.

If the reason is not compelling to a human editor, it probably is not compelling as a content strategy.

Part 8: Content value and evidence

Now ask the uncomfortable question:

Is this page worth retrieving?

Look for:

  • original examples
  • first-party data
  • case studies
  • technical observations
  • screenshots
  • methods
  • expert interpretation
  • primary-source citations
  • clear definitions
  • decision frameworks
  • and unique local/business information

Flag commodity pages that simply summarize what everyone else says. Google now explicitly recommends non-commodity, firsthand content for generative Search. Use that as permission to publish less generic material.

Part 9: Local entity consistency

For local businesses, compare:

  • website
  • Google Business Profile
  • Bing Places
  • structured data
  • major professional profiles
  • location pages
  • person pages
  • and tracking-number implementation

Review:

  • name
  • address
  • phone
  • hours
  • services
  • people
  • URLs
  • categories where relevant
  • service areas

The point is not ceremonial NAP perfection. The point is coherent identity.

Part 10: Crawler observability

If logs or Cloudflare are available, inspect:

  • known AI crawler requests
  • verification status
  • top paths
  • status codes
  • blocked requests
  • robots compliance
  • crawl waste
  • and frequency trends

Do not interpret crawl volume as visibility. Use it to validate access. A crawler cannot cite a page it cannot reach through the relevant path. It can also crawl a page thousands of times and never cite it.

Different stage.

Part 11: Platform visibility

Use first-party platform data where available.

Google Search Console

Generative AI visibility reports.

Bing Webmaster Tools

  • Citations.
  • Cited pages.
  • Grounding queries.
  • Topics/intents.
  • Citation share where available.

OpenAI

Referral traffic and crawler access.

Infrastructure

Crawler logs.

Do not replace this data with a third-party universal score when primary data exists.

Part 12: Business outcomes

Finally:

Does any of this matter commercially?

Measure:

  • AI referrals
  • landing pages
  • forms
  • calls
  • leads
  • revenue
  • assisted conversions
  • branded search
  • self-reported discovery source

This is where the audit reconnects with marketing. A technically perfect AI-search implementation with zero relevance to customer acquisition is an interesting engineering exercise. Mithril is a marketing company. Outcome matters.

Myth BustedA popular claim.
A closer look.

Follow the evidence

Myth: An AI audit is just SEO plus llms.txt

What the evidence says

A useful audit checks the information journey, from access to observable outcomes.

Technical SEO supplies much of the foundation. Crawler purposes, rendering differences, source clarity and platform-specific reporting add questions that deserve explicit checks.

How to prioritize findings

Do not give the client 187 red warnings.

Prioritize by:

  • Business impact
  • Dependency
  • Confidence
  • Effort
  • Breadth

Start with blockers.

Example:

OAI-SearchBot receives 403 on all service pages.

That is an access blocker if ChatGPT search visibility matters. Then structural problems.

Example:

Four duplicate location pages compete for the same office.

Then clarity/evidence improvements.

Example:

Attorney pages lack explicit practice-area relationships.

Then experimental enhancements.

Example:

Create a curated llms.txt resource for the documentation section.

This sequence matters. Do not polish machine-readable files while the canonical service page returns 500.

The four finding types

We like four types:

Blocker

Prevents or materially interferes with access/discovery.

Conflict

Creates contradictory or competing information.

Weak signal

Information exists but is ambiguous, thin, stale, or poorly connected.

Opportunity

A useful enhancement after foundations are correct.

This is more actionable than:

  • Critical
  • High
  • Medium
  • Low

based on whatever arbitrary scoring formula the audit vendor invented.

Mithril LabsTry it.
Observe it.
Learn from it.

Try it on your site

Example finding

Finding

Primary Phoenix office phone depends on CallRail DNI and is absent from the initial HTML.

Evidence

  • Initial response contains a placeholder number.
  • Rendered DOM contains a tracking number.
  • LocalBusiness schema contains the main corporate number.
  • Phoenix location page describes a separate direct office line.

Why it matters

Machines processing different page states receive three phone identities for the location.

Recommendation

Keep DNI for attribution. Establish one stable canonical Phoenix office number in first-party location data and structured data. Ensure the pre-script page state remains understandable and document tracking-number purpose.

Confidence

High. This is observable implementation behavior.

That is what a useful technical finding looks like.

It does not say:

“Phone schema missing, score -4.”

LumenWhat a tool can check

Where Lumen fits

The connection to Lumen is the action plan: each technical finding should have evidence, an owner and a useful next step. Page responses, links, metadata, rendering and structured data provide the website evidence.

Add available platform reports, request logs and sales feedback as separately identified sources. Explain how they support the recommendation and where further access or investigation is needed.

The final audit output

The deliverable should answer:

  • What is blocking machine access?
  • What information conflicts?
  • Which entities are unclear?
  • Which pages are weak retrieval candidates?
  • What content contributes unique value?
  • Which crawler policies are intentional?
  • What platform visibility exists today?
  • What should be fixed first?
  • What can be tested later?

Then translate that into a finite optimization plan. Not an endless warning list. That is the whole point.

Myth BustedA popular claim.
A closer look.

Follow the evidence

Myth: We need an AI readiness score

What the evidence says

Prioritized findings are more useful when your team needs to decide what to fix.

Record the issue, evidence, impact, owner and verification step. If a score is used, explain its inputs and limits so it can be interpreted alongside the findings.

The takeaway

Deliver findings with evidence, priority, ownership and a verification step. Start with access blockers and conflicting facts, then improve structure and content quality. Use experiments where the evidence is still uncertain.

Sources and primary references

  1. Google, Optimizing your website for generative AI featuresdevelopers.google.com
  2. Google, JavaScript SEOdevelopers.google.com
  3. Bing Webmaster Guidelinesbing.com
  4. Bing AI Performancebing.com
  5. OpenAI, Publishers and Developers FAQhelp.openai.com
  6. Cloudflare, AI Crawl Controldevelopers.cloudflare.com

The AI search guide

Where this fits

AI Search OptimizationThe starting point: how websites get crawled, retrieved, understood and cited.
  1. Access

    AI Crawlability
  2. Retrieval

    How AI Search Finds Sources
  3. Understanding and evidence

    Entities and Evidence
  4. Citation and outcome

    Measuring AI Search Visibility
Technical AI Search AuditThe capstone: the audit that tests every stage.
Cloudflare and AI PolicyTimely: Cloudflare’s controls as of September 18, 2026.