Free tool

AI crawler robots.txt generator

Decide, crawler by crawler, which AI companies may read your pages: for training, for search answers, or when a person asks. Then copy the result or merge it with your own robots.txt.

Training: collects content to train AI models
Search: builds the index behind AI search answers
User-initiated: fetches a page because a person asked
Paste it here. Rules for the crawlers above are replaced; everything else is kept as it is.

Your robots.txt


        

Then see who really reads you

robots.txt is a request. To see which robots actually fetch your pages, and which AI assistants send visitors, start free and open the AI & Search tab, or read how it works.

Most sites want something between "let every AI read everything" and "block them all". The usual middle path is to stay visible in AI search answers, where a link can send you visitors, and to opt out of having your content used to train models. This tool writes that file, or any other combination you prefer.

The three kinds of AI crawler

The crawlers fall into three groups, and the group tells you what blocking costs.

  • Training. These collect pages that may be used to train models. Blocking them is a request that your content is left out of future training. It does not remove what was collected already.
  • Search. These build the index behind an AI assistant's search answers. Blocking them keeps you out of those answers, and out of the visitors they send.
  • User-initiated. These fetch one page because a person asked an assistant about it. They are not crawling the web on their own.

Google-Extended and Applebot-Extended are different again: they are not crawlers but robots.txt switches. Googlebot and Applebot still crawl you for search; the switch only controls whether that content is used for AI.

How to use the result

Save the output as robots.txt in the root of your site, so it is served at https://your-site.com/robots.txt. Each subdomain needs its own file. Check it with the vendor's own tools where they offer one, and expect a delay: crawlers cache the file, and Meta says a change can take up to 24 hours to apply.

If you paste an existing file, rules that already name a crawler in the list are replaced, and everything else, your other user agents, Disallow lines and Sitemap entries, stays as it was.

Honest limits

robots.txt is a convention, not a lock. The well-known crawlers say they follow it, but nothing forces them to, and a robot that does not identify itself is not affected at all. Three vendors also say plainly that fetchers acting on a person's request may not follow it: OpenAI says robots.txt rules may not apply to ChatGPT-User, Perplexity says Perplexity-User generally ignores them, and Meta says Meta-ExternalFetcher may bypass them. To keep such a fetcher out you need to act on the request itself, for example in your CDN or firewall.

Blocking also has a price. A search crawler you block cannot cite you, and a user-initiated one you block cannot read your page when someone asks about it. Decide that page by page if you can, not as a reflex.

Where each crawler is documented

Vendors add and rename crawlers, so check these pages when you review your file. Some crawlers are missing here on purpose: a name goes on the list only when its operator documents it.

Questions

Which setting should I pick?
If you want visitors from AI answers but not training use, pick "Allow AI search, block training". If you publish for a living and want no AI use at all, pick "Block all AI", knowing that it also removes you from AI search.
Will blocking GPTBot remove my site from ChatGPT?
No. GPTBot is for training. ChatGPT's search feature uses OAI-SearchBot, and a person's request uses ChatGPT-User. OpenAI documents each separately.
Does blocking Google-Extended affect Google Search?
No. Google says Google-Extended does not affect inclusion in Google Search and is not a ranking signal. It controls use of your content for Gemini training and grounding.
Does this tool send my robots.txt anywhere?
No. It runs in your browser, and what you paste stays on this page.

The newsletter

Notes on measuring what matters.

Occasional emails from Albi: new posts, one chart worth reading, what changed in trckable.

We send one email to confirm your address, and nothing more until you confirm. Every newsletter has a link to leave. What we keep, and who sends it: privacy.

Counted, never watched.