Whether to allow AI training is a business decision about the material you publish and how you want it used. Search visibility and training permissions need separate consideration.
For service businesses publishing freely available educational material, our starting position is to consider allowing training access where there is no specific reason to restrict it. That is a business judgment, not a claim that training access improves rankings or citations.
Publishers, subscription businesses and organizations with licensed or sensitive material may reach a different conclusion. Identify the content, rights and commercial interests involved, then document a policy your team can maintain.
Start with the fact that search and training are different
OpenAI gives publishers separate controls for OAI-SearchBot and GPTBot. Google provides Google-Extended as a separate control for certain Gemini-related training and grounding uses while saying it does not affect Google Search inclusion or ranking.
Cloudflare increasingly distinguishes Search, Training, and Agent traffic.
Search
- The job
- Discovers and indexes pages for search results and AI answers.
- Examples
- Googlebot, including Google Search’s AI features. OAI-SearchBot for ChatGPT search.
- Starting policy
- Allow where discovery is desired.
Training
- The job
- Collects content that may be used to train or improve models.
- Examples
- GPTBot. The Google-Extended token for certain Gemini uses.
- Starting policy
- Make an explicit business decision.
Agent
- The job
- Acts for a person: reads, compares, checks availability, fills out a form.
- Examples
- Browser agents, and Cloudflare’s Agent category.
- Starting policy
- Allow where useful and safe, with normal security controls.
So the ecosystem itself is moving toward a useful idea:
Training permission should be a distinct business decision.
Now we can actually discuss it.
What are you trying to protect?
This is the first question. For a newspaper, research publisher, licensed database, paid analyst service, or company with proprietary documentation, the answer may be obvious. The public content itself can be a meaningful economic asset.
Training access may raise:
- licensing concerns,
- competitive concerns,
- contractual obligations,
- rights-management questions,
- or business-model questions.
Those are real. Now consider a local plumbing company.
Its website explains:
- water heater repair,
- drain cleaning,
- service areas,
- emergency availability,
- and why a toilet keeps running.
Identify the commercial interest you want the policy to protect. Public educational material, paid archives and licensed data may call for different decisions.
The service-business calculus is different
Many service businesses publish content because they want maximum understanding.
They want people to know:
- what they do,
- where they do it,
- who does the work,
- why they are qualified,
- what problems they solve,
- and how to contact them.
That is fundamentally different from selling access to information. If the primary business model is legal representation, roofing, accounting, dentistry, consulting, or marketing services, the article explaining the basics of the service is usually marketing material.
It exists to travel. That does not automatically mean every form of reuse should be permitted. It does mean “protect the content” needs to be weighed against why the content was published in the first place.
Our default position
For ordinary public service-business content, our starting position is generally:
- Allow major search access.
- Allow useful AI search access.
- Do not block training automatically.
- Restrict training when there is a real reason.
That reason might include:
- proprietary research,
- paid intellectual property,
- licensed third-party content,
- contractual restrictions,
- sensitive datasets,
- competitive product documentation,
- regulatory considerations,
- or a deliberate publisher-rights strategy.
The big caveat
We do not have evidence that allowing model training improves AI-search rankings or citation frequency. Treat a visibility benefit from training access as unproven. A training crawler and a search crawler may be entirely separate.
A model can also use retrieval systems that access current web information without that information having been part of model training. So our recommendation is strategic, not algorithmic.
We are saying:
For many service businesses, the cost of allowing training may be low enough that blocking it provides little practical benefit.
That is different from saying:
Allow GPTBot and ChatGPT will rank you higher.
No.
Observe it.
Learn from it.
Try it on your site
The “what are we protecting?” exercise
Take five public pages from the website.
For each page, ask:
- What information here is genuinely proprietary?
- Would a competitor gain something meaningful by learning it?
- Is the value in the words themselves, or in the service behind them?
- Is any content licensed from somebody else?
- Would reuse create a legal, contractual, or reputational concern?
- Is this page intentionally designed for broad public education?
Now classify the page:
- Open marketing content
- Strategically sensitive content
- Licensed/restricted content
- Private content that should not be public at all
This exercise often reveals that crawler policy should not be one global emotional reaction. Different content classes may deserve different treatment.
Publishers are different
A publication selling original reporting has a different economic relationship with its content. The article may be the product.
The reporting itself may have required:
- journalists,
- travel,
- records requests,
- interviews,
- photography,
- data analysis,
- legal review,
- and substantial cost.
A model learning from that material raises a different business question from a roofer publishing “5 Signs You Need a New Roof.” We should not flatten those businesses into one rule.
Good technical policy follows the business model.
What about law firms?
Law firms are interesting because their public content can be extensive and valuable. But the economic product is usually representation, not exclusive access to a general article about Arizona car accident claims.
Our default would still lean open for ordinary marketing content.
However, a firm may have:
- original research,
- proprietary settlement databases,
- internal templates,
- paid resources,
- or sensitive client material.
Those should be treated differently. And private material should be protected with actual authentication, not robots.txt.
A closer look.
Follow the evidence
Myth: “Blocking training makes your content safe from AI”
What the evidence says
A training restriction communicates a specific preference; its effect depends on the operator and enforcement.
Review what is publicly accessible, what controls the provider supports and whether sensitive material belongs behind authenticated access.
A closer look.
Follow the evidence
Myth: “Allowing training guarantees the model knows your brand”
What the evidence says
Training permission does not establish how a future model will represent or recommend your business.
Evaluate permissions on their own merits. Use observable search and referral data to assess discovery rather than treating training access as a placement promise.
So why lean open?
Because for many service businesses:
- the information is already intentionally public,
- the content is not itself the commercial product,
- the downside is limited,
- the business benefits when its concepts and expertise circulate,
- and maintaining highly restrictive crawler policies creates another technical surface that can accidentally block useful systems.
That is enough for us to default toward openness. It is not enough for us to tell every business to do the same thing.
The better policy: open by purpose, not by accident
We prefer:
- Public marketing content: broadly accessible unless there is a reason not to be.
- Commercially sensitive material: review.
- Licensed material: respect the license.
- Private content: protect it properly.
- AI search crawlers: allow where visibility is desired.
- Training crawlers: make an explicit business decision.
- Agents: allow useful interactions while enforcing normal security.
This is much more mature than flipping “Block AI Bots” because the setting exists.
Cloudflare’s new controls make this conversation easier
Cloudflare’s separation of Search, Training, and Agent traffic is valuable because it maps reasonably well to the actual business questions.
Do we want to be discovered?
Do we want models trained on this?
Do we want software agents interacting here?
Different question. Different switch.
Observe it.
Learn from it.
Try it on your site
Page-class policy
If you want more granularity, create content classes.
Class A: Public marketing and educational content
Default: broadly accessible.
Class B: High-value original research
Default: explicit review.
Class C: Licensed or partner content
Default: follow contractual rights.
Class D: Customer-only resources
Default: authenticated.
Class E: Internal/private material
Default: not publicly accessible.
Then map crawler policy to those classes. This is often more rational than pretending every URL on a domain has the same value or risk.
The strategic risk of over-blocking
The obvious risk of allowing access is unwanted reuse. The less obvious risk is blocking systems that support discovery. That is why search and training must remain separate in the policy.
If a business says:
“We do not want training.”
Fine. Then preserve search access where possible. OpenAI and Google both provide examples of more granular controls. Use them.
The strategic risk of underthinking it
The opposite mistake is saying:
“We are a marketing site, so who cares?”
You should still know:
- what is being crawled,
- by whom,
- for what purpose,
- and whether the content includes rights or data you are actually allowed to expose.
Openness should be a decision. Not neglect.
How Mithril would frame the client conversation
Use these questions to define the policy:
- Is your content itself a product?
- Do you license third-party material?
- Do you publish proprietary research?
- Do you want ChatGPT and similar systems to surface your public pages?
- Do you have a legal or contractual reason to restrict training?
- Do different site sections deserve different treatment?
Then configure the technical controls around the answers.
LumenWhat a tool can check
How Lumen fits
Compare the documented business policy with the technical settings. Keep separate records for search access, training permissions and agent behavior.
For example, a business may intentionally allow OAI-SearchBot while restricting GPTBot. The review should confirm the relevant directives and responses, then flag any conflict for the person responsible for the policy.
Our position in one sentence
Decide what you are protecting, separate training from search visibility and choose a policy that fits your business model. Record the reason for the choice so your team can review it when the content or controls change.
Sources and primary references
- OpenAI, Publishers and Developers FAQhelp.openai.com
- Google, Google-Extendeddevelopers.google.com
- Google, About crawlingdevelopers.google.com
- Cloudflare, AI Crawl Controldevelopers.cloudflare.com
- Cloudflare, managed robots.txt and Content Signalsdevelopers.cloudflare.com

