AI crawler policy

llms.txt Implementation and AI Crawler Policy That Decides Which AI Tools Can Read Your Website

llms.txt implementation publishes a curated map of your most important pages, and an AI crawler policy decides which AI tools may read your site. Because llms.txt is an emerging convention rather than a required standard, we set your robots.txt and firewall rules for each AI crawler first, then add llms.txt as an optional extra.

  • Published pricing from $799 per month
  • Facts checked against vendor documentation
  • Reply within 24 hours

Free, with no obligation

Get a free AI search review

Tell us where to reply. A strategist looks at how AI tools describe your business and replies within 24 hours.

This form needs JavaScript. Call +1 (510) 399-5523, email jd@aisearchrankings.com, or use the contact page.

We use your details only to reply. We never sell your information. Read our privacy policy.

Prefer to talk now? Call +1 (510) 399-5523.

The same work supports visibility across these AI search experiences

  • ChatGPT logoChatGPT
  • Google AI Overviews logoGoogle AI Overviews
  • Google AI Mode logoGoogle AI Mode
  • Perplexity logoPerplexity
  • Microsoft Copilot logoMicrosoft Copilot
  • Grok logoGrok

The platform mix and monitoring coverage are confirmed for your market during onboarding.

Key takeaways

  • llms.txt is an emerging convention rather than a required standard, and Google states no AI text file is needed for its AI features.
  • OpenAI, Anthropic, and Google each separate search access from model training, so a deliberate crawler policy matters more than the file.
  • A crawler with its own robots.txt group ignores your general rules, so private folders must be blocked again inside that group.
  • We include llms.txt inside the monthly package rather than selling it as a standalone service.
Crawler policy first, llms.txt second
  1. Decide per crawlerSearch access and model training are separate choices for each vendor.
  2. Set robots.txtRules that match your decisions and still protect private folders.
  3. Match your firewallBot protection that allows what robots.txt allows.
  4. Publish llms.txtAn optional, curated list of your most important pages.

llms.txt is an emerging convention, not a required standard. The crawler policy is the part that matters today.

Three files compared

robots.txt controls access. llms.txt only offers a map.

Three plain text files at the root of a website are often confused. Only one of them can allow or block a crawler, and only two are established standards.

  • robots.txt

    Internet standard, RFC 9309

    Who reads it
    Search engines and AI crawlers that follow the Robots Exclusion Protocol.
    What it does
    Allows or blocks paths for each crawler you name.
  • sitemap.xml

    Established protocol

    Who reads it
    Google, Bing, and other search engines.
    What it does
    Lists the addresses you want discovered, with optional update dates.
  • llms.txt

    Proposal, not a standard

    Who reads it
    Tools that choose to look for it.
    What it does
    Summarizes your site and links your most important pages. It cannot block anything.
Comparison of three plain text files published at the root of a website, as of September 13, 2026.
  • Google does not require it

    Google states that no new machine readable file, AI text file, or special markup is needed to appear in AI Overviews or AI Mode.

  • It cannot block a crawler

    Access decisions belong in robots.txt and your firewall. llms.txt has no allow or disallow function.

Sources: Google Search Central: AI features and your website, RFC 9309: Robots Exclusion Protocol, the llms.txt proposal.

Want this done for your business?

Get a free review of where you stand today. A strategist replies within 24 hours.

A real llms.txt

What a useful llms.txt looks like, taken from our own website

An excerpt from the llms.txt file AI Search Rankings publishes. It describes the company in plain words and links the pages worth citing.

Excerpt from /llms.txt

# AI Search Rankings - llms.txt

## What AI Search Rankings Is

AI Search Rankings is an Answer Engine Optimization (AEO) agency that helps
businesses become easier for AI search engines to find, understand, trust,
cite, and recommend.

## Verified Company Facts

- Brand name: AI Search Rankings
- Category: Answer Engine Optimization (AEO) agency
- Founder: Jagdeep Singh, Founder and Lead Answer Engine Optimization Strategist
- Founded: 2025
- Service area: United States

## Key Pages for Citation

- Brand facts: https://www.aisearchrankings.com/brand-facts.php
- Pricing: https://www.aisearchrankings.com/pricing.php
- Services hub: https://www.aisearchrankings.com/services/

A shortened excerpt. Open the full llms.txt file.
  • A plain description

    One short paragraph that says what the business is, in the words you want repeated.

  • Facts stated exactly

    Name, category, founder, founding year, and service area, matching your brand facts and your pages.

  • Links worth citing

    The pages you most want an AI tool to read, each with a clear label.

  • Kept in sync

    A stale llms.txt repeats old facts, so it is updated whenever your pages or policies change.

See how AI tools describe your business

We check your most important buyer questions for free and show you the first fixes to make.

How we implement

Crawler policy first, llms.txt second

The decisions that affect visibility come first. The optional file comes last and stays in sync with your facts.

  1. Decide per crawlerSearch access and model training are separate decisions for each vendor.
  2. Set robots.txtRules that match your decisions and still protect private folders in every group.
  3. Match your firewallBot protection that allows what robots.txt allows, so rules are not silently overridden.
  4. Publish llms.txtAn optional, curated list of your most important pages and exact facts.
  5. Recheck monthlyVendors add and rename crawlers, so the policy is reviewed every month.

Related work: technical SEO services and ChatGPT optimization services.

Not sure where to start?

Tell us about your business, and we will point you to the few changes that matter most.

Checked against vendor documentation

Which AI crawlers to allow, and what each one controls

Each AI vendor publishes separate controls for search access and for model training. Checking them is the cheapest, fastest fix in AI search, and one wrong rule can hide your pages without any warning on your website.

Your websitePages, facts, and answers
robots.txt and firewallDecide which crawlers get in

Model training controls

  • Google-Extended
  • GPTBot
  • ClaudeBot

Allow or block these as a business decision. Each vendor documents them separately from its search crawler.

How robots.txt and firewall rules decide which AI crawlers can read your website. Crawler names come from vendor documentation, checked September 13, 2026.
Search access and training controls by AI vendor
AI productSearch accessTraining controlWhat the training control changes
Google AI Overviews and AI ModeGooglebotGoogle-ExtendedGoogle states that Google-Extended does not affect inclusion in Google Search and is not a ranking signal in Google Search.
Microsoft CopilotbingbotNone listedThe cited documentation lists no separate training crawler.
ChatGPTOAI-SearchBotGPTBotBlocking it tells OpenAI not to use your content for model training. Appearance in ChatGPT search is controlled separately by OAI-SearchBot.
PerplexityPerplexityBotNone listedThe cited documentation lists no separate training crawler.
ClaudeClaude-SearchBotClaudeBotAnthropic states that blocking it excludes your future content from its model training datasets.

What the documentation says

  • Google states that a page must be indexed and eligible to show with a snippet in Google Search to appear as a supporting link in AI Overviews or AI Mode.
  • Google states that no special schema, AI text file, or new machine readable file is required to appear in these features.
  • The nosnippet, data-nosnippet, max-snippet, and noindex controls limit what Google can show, including in AI features.
  • Clicks and impressions from AI features are included in the Search Console Performance report under the Web search type.
  • Bing Webmaster Tools added an AI Performance report in public preview in February 2026. It shows how often your pages are cited in Microsoft Copilot and AI generated summaries in Bing.
  • Bing supports IndexNow, which notifies Bing when a page is added, changed, or removed.
  • OpenAI states that a robots.txt change can take about 24 hours to take effect.
  • Perplexity publishes the IP ranges for both agents, so a server firewall can allow them without allowing everyone.
  • Anthropic states that its crawlers honor robots.txt, including the Crawl-delay extension.

The rule most people miss. Under the robots.txt standard, a crawler obeys the most specific group that names it. If you add a group for one crawler, that crawler stops reading your "User-agent: *" rules, so any private folders you block there must be blocked again inside the new group.

Stay visible in AI search and opt out of model training

# Search crawlers (Googlebot, bingbot, OAI-SearchBot,
# PerplexityBot, Claude-SearchBot) need no new group.
# They follow your "User-agent: *" rules unless a group
# names them. Remove any rule that blocks them.

# Opt out of OpenAI model training.
# ChatGPT search is controlled by OAI-SearchBot instead.
User-agent: GPTBot
Disallow: /

# Opt out of Anthropic model training.
User-agent: ClaudeBot
Disallow: /

# Opt out of Gemini training and Gemini Apps grounding.
# Google states this does not affect Google Search.
User-agent: Google-Extended
Disallow: /

Test changes in Google Search Console and Bing Webmaster Tools after you publish them.

Sources, checked September 13, 2026: Google Search Central: AI features and your website, Google common crawlers: Google-Extended, Microsoft Learn: web search in Microsoft Copilot, Bing Webmaster Blog: AI Performance public preview, OpenAI crawler documentation, Perplexity crawler documentation, Anthropic crawler documentation, RFC 9309: Robots Exclusion Protocol.

Not sure what your website allows AI crawlers to do?

Ask for a free review, and we will look at your robots.txt file and tell you whether it blocks the crawlers listed above.

Free self check, about 2 minutes

AI crawler policy and llms.txt readiness check

Answer eight questions about which AI systems can read your website. The check puts the proven work first and treats llms.txt honestly as optional.

Choose Yes, No, or Not sure for each question, then select See my score and fixes. Open the details under a question to see why it matters, how to check it, and the fix.

  1. Have you decided, crawler by crawler, which AI search and training crawlers may read your site?
    Why it matters and how to check

    Why it matters. OpenAI, Anthropic, and Google each separate search access from model training, so one blanket rule rarely matches what a business wants.

    How to check. List OAI-SearchBot, GPTBot, Claude-SearchBot, ClaudeBot, PerplexityBot, and Google-Extended, and write your decision for each.

    The fix. Record a decision for each crawler and set robots.txt to match it.

  2. Has someone reviewed your robots.txt file in the last 90 days?
    Why it matters and how to check

    Why it matters. AI vendors add and change crawlers, so an old file can block or allow things nobody chose.

    How to check. Open your website address followed by /robots.txt and check each rule against your written decisions.

    The fix. Review robots.txt every quarter and after every vendor documentation change.

  3. Does every crawler specific group repeat the Disallow rules that protect private folders?
    Why it matters and how to check

    Why it matters. Under the robots.txt standard, a crawler obeys only the most specific group that names it, so it stops reading your general rules.

    How to check. Compare each named crawler group in robots.txt with your "User-agent: *" group.

    The fix. Copy every private Disallow line into each crawler specific group.

  4. Does your firewall or bot protection match your robots.txt decisions?
    Why it matters and how to check

    Why it matters. A firewall can block a crawler your robots.txt allows, and the block is invisible on the website.

    How to check. Check your CDN or security settings for blocked AI crawler categories.

    The fix. Align firewall rules with your crawler decisions using the vendors' published IP ranges.

  5. Does one page hold your governed business facts in plain text?
    Why it matters and how to check

    Why it matters. A single facts page gives every crawler and every human the same accurate description.

    How to check. Check for a page that lists your name, services, area, contact details, and founder.

    The fix. Publish a plain text facts page and keep it current.

  6. Is your XML sitemap current and referenced in robots.txt?
    Why it matters and how to check

    Why it matters. A current sitemap is the established way to tell crawlers which pages exist.

    How to check. Look for a Sitemap line in robots.txt and open the file it points to.

    The fix. Reference an automatically updated sitemap in robots.txt.

  7. Does your site have an llms.txt file listing your most important pages?
    Why it matters and how to check

    Why it matters. llms.txt is an emerging convention rather than a required standard, and Google states no AI text file is needed for its AI features, so treat it as low cost and optional.

    How to check. Open your website address followed by /llms.txt.

    The fix. Publish a short llms.txt file only after the higher weight items are complete.

  8. Is your crawler policy checked whenever you launch, move, or retire pages?
    Why it matters and how to check

    Why it matters. Site changes can quietly expose private areas or hide new public ones.

    How to check. Check whether your launch checklist includes robots.txt, llms.txt, and the sitemap.

    The fix. Add crawler policy checks to every launch and redesign checklist.

Proof with its limits attached

Documented results, labeled honestly

Each result states what was measured, where the data came from, and what it does not prove. Past results show what the approach can do. They do not guarantee your outcome.

  • The Toolsforhuman.com website shown on a desktop monitor

    Measured on Microsoft Copilot

    Verified result

    26,300Microsoft Copilot citations

    Toolsforhuman.com Microsoft Copilot

    AI Search Rankings used schema-first Answer Engine Optimization, layered JSON-LD, robust canonicals, and multilingual hreflang to help Toolsforhuman.com earn 26,300 Microsoft Copilot citations across 160 plus pages, including Spanish-language pages, in a sample window of about three months.

    Limitation. The result applies to the documented sample window and does not guarantee future citations.

  • Laser Tagging Inc. website hero section before and after the redesign

    Measured on Google AI Overviews and AI Mode

    Client-reported private Google Search Console result

    4,460Google generative AI impressions

    Laser Tagging Inc. Google AI Overviews and AI Mode

    AI Search Rankings rebuilt the Laser Tagging Inc. digital brand and website, expanded its service-area relevance across Newark, Fremont, San Jose, and nearby Bay Area communities, and documented 4,460 client-reported Google generative AI impressions in the selected last-three-month window.

    Limitation. The report measures Google generative AI impressions. It does not prove a ChatGPT recommendation, a click, a booking, a lead, or revenue.

    Newer data shared by the client: a Google Search performance export for search type Web shows 6,109 impressions in the last 3 months as of September 13, 2026, compared with 947 in the previous 3 months. That total covers all Google web results, not only AI features. Checked September 13, 2026. Open the search performance export shared by Laser Tagging Inc. (opens in a new tab)

How AI Search Rankings labels and measures results

Start with your own baseline

Every engagement starts with a measured baseline. Ask for a free review and see where your business stands today.

Why AI Search Rankings

Why businesses choose us for llms.txt Implementation

Senior expertise, published prices, and honest reporting from the first conversation. This is what partnering with us looks like.

  • Led by a named expert

    Jagdeep Singh, Founder and Lead Answer Engine Optimization Strategist, brings 12+ years of SEO and digital marketing experience and 9 Google Skills certifications you can verify.

  • Prices published before you call

    Three monthly packages at $799, $1,500, and $3,500, each with a written list of the work delivered every month.

  • Reports you can check yourself

    We record the real AI answers for your buyer questions and list the work shipped beside them, never invented rankings.

  • Proof with its limits stated

    Results such as 26,300 Microsoft Copilot citations for Toolsforhuman.com are published with how they were measured and what they do not prove.

  • Honest about what no one controls

    We guarantee the work, not rankings or citations, so you are never sold a promise that cannot be kept.

  • A fast, straight answer

    A strategist replies within 24 hours, Monday to Friday, and tells you plainly when we are not the right fit.

Ready to work with a team that shows its work?

Start with a free review and a clear, published scope, with no obligation.

Published pricing

What llms.txt Implementation costs, and which scope fits you

Every price is a fixed monthly scope of work. It is never a promise about rankings or AI citations, because other companies control those systems.

llms.txt Implementation is included from Local AI Answer Authority at $3,500 per month, and in every higher package.

Local or service-area business

Three published monthly packages, built around local and service-area businesses. Each one lists exactly what is delivered every month.

  • Local small businesses

    Local AI Answer Starter

    $799 per month

    Help nearby customers find, trust, and choose your business across Google, Google Maps, and AI answers.

    Local buyer prompts tracked
    25
    Local competitors compared
    3
    Page optimized monthly
    1
    New local answers monthly
    3 to 5
    See everything in Local AI Answer Starter
  • Competitive & multi-location businesses

    Local AI Answer Authority

    $3,500 per month

    Build a complete local authority system across Google, Maps, AI answers, reviews, and content.

    Local buyer prompts tracked
    150 to 250
    Local competitors compared
    10
    Pages optimized monthly
    5 to 8
    New local answers monthly
    10 to 20
    See everything in Local AI Answer Authority

Not sure which package fits? Answer four questions in the AI Search Fit Check, or compare every package line by line.

National, B2B, SaaS, or online business

Our three published packages were built around local and service-area businesses, which is why their names say Local. National, B2B, SaaS, and online businesses are also served, and the work itself is the same: buyer questions tracked, pages improved, facts corrected, and results reported.

What changes is what the questions and pages cover, such as products, categories, comparisons, and several markets. Start with the free package fit review. We compare your buyer questions, pages, and competitors against the published scopes and confirm in writing whether one of them fits before any work starts.

Request a free package fit review
Estimate your monthly scope

Count services, comparison questions, and cost questions buyers ask.

Your most valuable service, product, or category pages.

The businesses buyers compare you with most often.

The largest published scope covers up to 250 buyer questions, 8 pages improved each month, and 10 competitors. The estimate is a starting point, and the fit review confirms it.

Straight answers

llms.txt and AI crawler policy questions, answered

Plain answers to the questions buyers ask most. If yours is not here, ask a strategist directly.

What is llms.txt?

llms.txt is a proposed convention for a plain text file, written in Markdown and placed at the root of a website, that summarizes the site and links its most important pages for large language models. It was proposed in 2024 and is not an official web standard.

Is llms.txt the same as robots.txt?

No. robots.txt is the Robots Exclusion Protocol, published as RFC 9309, and it tells crawlers which paths they may access. llms.txt does not allow or block anything. It only offers a curated summary and a list of links, so access decisions still belong in robots.txt and your firewall.

Does Google use llms.txt for AI Overviews?

Google states that you do not need to create new machine readable files, AI text files, or special markup to appear in AI Overviews or AI Mode. Indexable pages that follow standard SEO practices are what its AI features rely on.

Which AI crawlers should we allow?

Most businesses allow the crawlers that power AI search answers, such as OAI-SearchBot, PerplexityBot, and Claude-SearchBot, and decide separately about training crawlers such as GPTBot, ClaudeBot, and Google-Extended. Each vendor documents its search and training crawlers separately, so the two decisions can differ.

Will blocking AI training crawlers remove us from AI answers?

Not for the vendors whose documentation we checked. OpenAI documents GPTBot separately from OAI-SearchBot, Google states Google-Extended does not affect Google Search, and Anthropic documents Claude-SearchBot separately from ClaudeBot. Blocking a search crawler is what removes pages from that tool's search answers.

What goes in a good llms.txt file?

A short description of the business, the facts you want repeated exactly, and links to your most important pages with a one line summary of each. It should match your brand facts and your visible pages, and it should be updated whenever those change.

Can llms.txt keep our content out of AI training?

No. llms.txt has no blocking function. Opting out of training is done with robots.txt rules for the training crawlers each vendor documents, backed by firewall rules if you need enforcement.

Why can new AI crawler rules open private folders?

Under RFC 9309, a crawler follows only the group that names it and ignores the rules written for all crawlers. A new group for an AI crawler must repeat every private Disallow line, or that crawler is allowed into those folders.

Does AI Search Rankings publish its own llms.txt?

Yes. AI Search Rankings publishes an llms.txt file at the root of its website that describes the company and links its key pages. An excerpt is shown on this page, and the full file is linked beside it.

How much does llms.txt implementation cost?

Crawler policy and llms.txt are included in every package at $799, $1,500, or $3,500 per month. We do not sell llms.txt as a standalone service, because the crawler policy is the part that affects visibility today.

Free, with no obligation

Make a deliberate AI crawler decision

Choose what you need help with and tell us where to reply. We review your robots.txt, firewall rules, and key pages and respond within 24 hours.

What happens after you send it

  1. We review your website and the answers AI tools give for your buyer questions.
  2. A strategist replies within 24 hours, Monday to Friday, 9am to 6pm Pacific Time.
  3. You get a clear recommendation, including when we are not the right fit.

Prefer to talk now?

This request form needs JavaScript. You can still reach a strategist by phone, text, or email using the links beside this form, or use the contact page.

  1. Your goal
  2. Where to reply
What do you want help with? Choose any that apply
Which best describes your business? Required

For example, yourbusiness.com. We look at it before we reply.

Where should we reply?

For example, your main service, the cities you serve, or a deadline.

We use your details only to reply to this request. There is no obligation, and we never sell your information. Read our privacy policy.