A technical AI-search audit checks whether important business information is accessible, clear, consistent and supported. It produces findings your team can investigate and act on.
Our approach follows seven observable stages:
- 1Access
Breaks when a crawler is blocked.
The system needs a route to the information before anything else can happen.
- 2Retrieval
Breaks when the page is a weak candidate.
The page needs to match the information the system is looking for.
- 3Understanding
Breaks when the entity is ambiguous.
Names, entities and relationships need to be clear and consistent.
- 4Evidence
Weakens when claims lack useful evidence.
Original experience and verifiable facts give the content a reason to be used.
- 5Selection
Breaks when a stronger source is chosen.
The system compares candidate sources for the particular question.
- 6Citation
A visible reference helps readers identify the source.
Measure visible citations and referral visits separately.
- 7Outcome
Breaks when visibility never becomes a lead.
Connect visibility with qualified inquiries and business results.
We inspect the website, its responses and the available reporting data. The result is a prioritized set of issues, each with evidence, an explanation and a recommended next step.
Part 1: Access
First question:
Can relevant machines reach the information?
We review:
- robots.txt
- crawler-specific directives
- meta robots
- X-Robots-Tag
- CDN/WAF controls
- bot-management policies
- rate limits
- authentication
- status codes
- redirects
- server responses
- and known crawler behavior where logs are available
We distinguish:
- search crawlers,
- training crawlers,
- agents,
- and other automated traffic.
Do not report:
“AI bots blocked”
without saying which bot, why it matters, and whether the state is intentional.
Observe it.
Learn from it.
Try it on your site
Access contradiction check
Compare:
- robots.txt policy
- Cloudflare/WAF policy
- page-level robots directives
- server response
If robots says Allow and the WAF returns 403, the infrastructure wins. If robots blocks the page, a crawler may never read the noindex on that page. If an AI-search crawler is blocked while visibility in that product is a business goal, document the mismatch.
The finding should connect technical state to intent.
Part 2: Raw HTML versus rendered page
Next:
What does the server actually send?
Compare:
- initial HTML
- rendered DOM
- interaction-dependent content
- third-party widgets
Important checks:
- Does the primary service exist in raw or reliably rendered content?
- Are important links crawlable?
- Are contact details stable?
- Does DNI replace phone numbers?
- Are location details injected?
- Is schema client-side?
- Do FAQ or tab contents require interaction?
- Does navigation exist before scripts initialize?
This is one of the strongest places Lumen can contribute. A browser screenshot cannot answer these questions.
Part 3: Discovery and site architecture
Important pages need paths.
Review:
- internal links
- navigation
- breadcrumbs
- sitemaps
- orphan pages
- crawl depth
- redirecting internal links
- parameter traps
- faceted navigation
- pagination
- and broken links
Then map page roles.
Which URL is the pillar?
Which URL owns the service?
Which URL represents the office?
Which URL represents the person?
Which supporting pages deepen the topic?
A website can have thousands of indexable pages and still have no coherent architecture.
Part 4: Canonical truth
Find competing versions.
Review:
- rel=canonical
- redirects
- sitemap URLs
- HTTP/HTTPS
- www/non-www
- parameters
- old slugs
- campaign copies
- staging copies
- duplicate content
- near-duplicate content
- and conflicting internal links
Then review facts.
Does one obvious source own:
- business name
- phone
- address
- hours
- people
- services
- pricing
- policies
- credentials
AI search makes stale canonical truth more visible because generative answers can surface the wrong version confidently. Give systems fewer wrong versions.
Part 5: Entity clarity
Inventory the things that matter.
- Organization
- People
- Locations
- Services
- Products/tools
- Credentials
- Key content assets
For each entity, ask:
- Is there a canonical page?
- Is the name consistent?
- Are relationships explicit?
- Does structured data reflect the page?
- Do internal links connect related entities?
- Do major external profiles agree?
This is especially important for local and professional-service businesses.
Part 6: Structured data
We do not score schema quantity.
We evaluate:
- validity
- appropriateness
- visible-content alignment
- entity relationships
- canonical identifiers
- duplicate/conflicting nodes
- JavaScript dependencies
- location accuracy
- person accuracy
- and eligibility for relevant search features
Remember:
Google says there is no special schema required for generative Search.
Schema is an encoding layer. Use it accurately.
Part 7: Retrieval readiness
Now ask whether important pages are useful retrieval candidates.
- Does each page have a clear topic?
- Does it answer the main information need?
- Are important facts explicit?
- Can the page stand alone?
- Does it contain relevant evidence?
- Is critical information surfaced reasonably early?
- Is the content current?
- Does the URL have a unique reason to exist?
- Are there stronger duplicate pages?
Bing’s current guidance gives us useful observable principles here:
- focused URLs,
- explicit facts,
- clear headings,
- current information,
- crawlable links,
- and independent verifiability.
These do not guarantee a citation. They create better source candidates.
Observe it.
Learn from it.
Try it on your site
The “why this page?” test
For each priority URL, finish:
“This page should be retrieved because...”
Weak
It targets the keyword.
Better
It is the canonical page explaining Mithril’s technical AI-search audit process, includes the full methodology, links to supporting diagnostic guides, and contains information not repeated elsewhere.
If the reason is not compelling to a human editor, it probably is not compelling as a content strategy.
Part 8: Content value and evidence
Now ask the uncomfortable question:
Is this page worth retrieving?
Look for:
- original examples
- first-party data
- case studies
- technical observations
- screenshots
- methods
- expert interpretation
- primary-source citations
- clear definitions
- decision frameworks
- and unique local/business information
Flag commodity pages that simply summarize what everyone else says. Google now explicitly recommends non-commodity, firsthand content for generative Search. Use that as permission to publish less generic material.
Part 9: Local entity consistency
For local businesses, compare:
- website
- Google Business Profile
- Bing Places
- structured data
- major professional profiles
- location pages
- person pages
- and tracking-number implementation
Review:
- name
- address
- phone
- hours
- services
- people
- URLs
- categories where relevant
- service areas
The point is not ceremonial NAP perfection. The point is coherent identity.
Part 10: Crawler observability
If logs or Cloudflare are available, inspect:
- known AI crawler requests
- verification status
- top paths
- status codes
- blocked requests
- robots compliance
- crawl waste
- and frequency trends
Do not interpret crawl volume as visibility. Use it to validate access. A crawler cannot cite a page it cannot reach through the relevant path. It can also crawl a page thousands of times and never cite it.
Different stage.
Part 11: Platform visibility
Use first-party platform data where available.
Google Search Console
Generative AI visibility reports.
Bing Webmaster Tools
- Citations.
- Cited pages.
- Grounding queries.
- Topics/intents.
- Citation share where available.
OpenAI
Referral traffic and crawler access.
Infrastructure
Crawler logs.
Do not replace this data with a third-party universal score when primary data exists.
Part 12: Business outcomes
Finally:
Does any of this matter commercially?
Measure:
- AI referrals
- landing pages
- forms
- calls
- leads
- revenue
- assisted conversions
- branded search
- self-reported discovery source
This is where the audit reconnects with marketing. A technically perfect AI-search implementation with zero relevance to customer acquisition is an interesting engineering exercise. Mithril is a marketing company. Outcome matters.
A closer look.
Follow the evidence
Myth: “An AI audit is just SEO plus llms.txt”
What the evidence says
A useful audit checks the information journey, from access to observable outcomes.
Technical SEO supplies much of the foundation. Crawler purposes, rendering differences, source clarity and platform-specific reporting add questions that deserve explicit checks.
How to prioritize findings
Do not give the client 187 red warnings.
Prioritize by:
- Business impact
- Dependency
- Confidence
- Effort
- Breadth
Start with blockers.
Example:
OAI-SearchBot receives 403 on all service pages.
That is an access blocker if ChatGPT search visibility matters. Then structural problems.
Example:
Four duplicate location pages compete for the same office.
Then clarity/evidence improvements.
Example:
Attorney pages lack explicit practice-area relationships.
Then experimental enhancements.
Example:
Create a curated llms.txt resource for the documentation section.
This sequence matters. Do not polish machine-readable files while the canonical service page returns 500.
The four finding types
We like four types:
Blocker
Prevents or materially interferes with access/discovery.
Conflict
Creates contradictory or competing information.
Weak signal
Information exists but is ambiguous, thin, stale, or poorly connected.
Opportunity
A useful enhancement after foundations are correct.
This is more actionable than:
- Critical
- High
- Medium
- Low
based on whatever arbitrary scoring formula the audit vendor invented.
Observe it.
Learn from it.
Try it on your site
Example finding
Finding
Primary Phoenix office phone depends on CallRail DNI and is absent from the initial HTML.
Evidence
- Initial response contains a placeholder number.
- Rendered DOM contains a tracking number.
- LocalBusiness schema contains the main corporate number.
- Phoenix location page describes a separate direct office line.
Why it matters
Machines processing different page states receive three phone identities for the location.
Recommendation
Keep DNI for attribution. Establish one stable canonical Phoenix office number in first-party location data and structured data. Ensure the pre-script page state remains understandable and document tracking-number purpose.
Confidence
High. This is observable implementation behavior.
That is what a useful technical finding looks like.
It does not say:
“Phone schema missing, score -4.”
LumenWhat a tool can check
Where Lumen fits
The connection to Lumen is the action plan: each technical finding should have evidence, an owner and a useful next step. Page responses, links, metadata, rendering and structured data provide the website evidence.
Add available platform reports, request logs and sales feedback as separately identified sources. Explain how they support the recommendation and where further access or investigation is needed.
The final audit output
The deliverable should answer:
- What is blocking machine access?
- What information conflicts?
- Which entities are unclear?
- Which pages are weak retrieval candidates?
- What content contributes unique value?
- Which crawler policies are intentional?
- What platform visibility exists today?
- What should be fixed first?
- What can be tested later?
Then translate that into a finite optimization plan. Not an endless warning list. That is the whole point.
A closer look.
Follow the evidence
Myth: “We need an AI readiness score”
What the evidence says
Prioritized findings are more useful when your team needs to decide what to fix.
Record the issue, evidence, impact, owner and verification step. If a score is used, explain its inputs and limits so it can be interpreted alongside the findings.
The takeaway
Deliver findings with evidence, priority, ownership and a verification step. Start with access blockers and conflicting facts, then improve structure and content quality. Use experiments where the evidence is still uncertain.
Sources and primary references
- Google, Optimizing your website for generative AI featuresdevelopers.google.com
- Google, JavaScript SEOdevelopers.google.com
- Bing Webmaster Guidelinesbing.com
- Bing AI Performancebing.com
- OpenAI, Publishers and Developers FAQhelp.openai.com
- Cloudflare, AI Crawl Controldevelopers.cloudflare.com

