Key points
Robots.txt is a text file placed at the root of a website. It tells crawlers which URLs they may or may not crawl. It controls crawling, not indexing: a blocked page can still appear in Google. It is also the file that opens or closes your site to AI crawlers.
- Four directives matter for Google: User-agent, Disallow, Allow and Sitemap.
- To remove a page from Google, use noindex, never robots.txt.
- Blocking GPTBot does not have the same effect as blocking OAI-SearchBot: one targets training, the other citation in ChatGPT.
What is a robots.txt file?
The robots.txt file applies the Robots Exclusion Protocol, created in 1994. This protocol became an IETF standard in September 2022: RFC 9309, co-written by Google engineers. A well-behaved crawler reads this file before crawling a site, then follows its rules.
According to Google, it is used to manage crawler traffic to your site (Introduction to robots.txt, December 10, 2025). It prevents your server from being overloaded and stops crawlers from wasting time on pages of no interest. It is neither a security tool nor a way to remove a page from Google.
Where is a site’s robots.txt located?
Always at the root of the host: https://www.exemple.fr/robots.txt. Type this address to read the file of any site. It only applies to the host, protocol and port where it is published. blog.exemple.fr or the http version need their own file.
What Google does depending on the server response
| File response | Google’s behavior |
|---|---|
| 200 | Google reads and applies the rules. It generally caches the file for up to 24 hours. |
| 404 or other 4xx (except 429) | Google assumes there are no restrictions and crawls everything. |
| 5xx or server unreachable | Google stops crawling for 12 hours, then relies on the last cached version for 30 days while retrying. |
| File larger than 500 KiB | Anything beyond the limit is ignored. |
These rules come from the robots.txt specification published by Google (updated on August 31, 2026). A robots.txt that returns a server error is therefore more dangerous than a missing file.
Which directives does Google understand?
| Directive | Role | Example |
|---|---|---|
User-agent | Names the targeted crawler; * targets all crawlers that have no dedicated group | User-agent: Googlebot |
Disallow | Forbids crawling of a path | Disallow: /panier/ |
Allow | Allows a path inside a forbidden area | Allow: /panier/aide/ |
Sitemap | Declares the absolute URL of a sitemap, outside any group | Sitemap: https://www.exemple.fr/sitemap.xml |
*stands for any sequence of characters and$marks the end of the URL:Disallow: /*.pdf$blocks PDFs.- Paths are case-sensitive:
/Admin/and/admin/are two separate rules. - In case of conflict, Google applies the most specific rule (the longest path); if they are equal, the least restrictive one.
- Google ignores
Crawl-delay. Other crawlers, such as Anthropic’s, take it into account.
# Annotated example
User-agent: *
Disallow: /recherche/
Disallow: /*?tri=
Allow: /recherche/aide/
Sitemap: https://www.exemple.fr/sitemap.xml
Robots.txt or noindex: which one should you use?
This is the most common mistake. Robots.txt stops the crawler from reading the page; it does not stop it from knowing its address. If other sites link to it, Google can show the URL without a description (Google introduction cited above).
| Goal | Right tool | Why |
|---|---|---|
| Keep low-value areas from being crawled (sorting, internal search, cart) | robots.txt | Saves crawling, especially on a large site |
| Remove a page from Google’s index | noindex tag or X-Robots-Tag header | Google must be able to crawl the page to read the instruction |
| Protect a staging site or a customer area | Password (server authentication) | Robots.txt is public and some crawlers ignore it |
Warning
Never combine noindex and Disallow on the same page. Once blocked by robots.txt, the page is no longer read. Google does not see the noindex and the URL can stay indexed. First remove the page from the index with noindex, and only then block crawling if needed. More details in our article on Google indexing.
On a large site, Google recommends robots.txt to permanently block URLs that should never be crawled. This is the topic of crawl budget.
How do you manage AI crawler access in robots.txt?
AI companies now publish several crawlers, each with a distinct role. Blocking them all at once often means dropping out of ChatGPT or Claude answers without meaning to. There are three families to tell apart.
| Family | Crawlers (as of September 24, 2026) | Effect of blocking |
|---|---|---|
| Model training | GPTBot (OpenAI), ClaudeBot (Anthropic), Google-Extended token | Your future content is left out of training data |
| Answer engine index | OAI-SearchBot (ChatGPT), Claude-SearchBot, Googlebot (which also feeds AI Overviews and AI Mode) | Your pages can no longer be cited as a source |
| On-demand reading for a user | ChatGPT-User, Claude-User | The assistant can no longer read the page on request; OpenAI states that ChatGPT-User may not follow these rules |
Three details taken from the companies’ documentation. At OpenAI, OAI-SearchBot makes sites appear in ChatGPT search. Robots.txt rules, however, may not apply to ChatGPT-User, whose visits are triggered by a user (OpenAI, crawler documentation). Anthropic describes its three crawlers, respects robots.txt and accepts Crawl-delay (Anthropic, help center). Finally, Google-Extended is not a crawler but a token. It controls the use of your content for Gemini, with no effect on your presence or ranking in Google (Google, list of crawlers, July 14, 2026).
A common trade-off: stay citable by answer engines without feeding training.
# Refuse training, stay citable
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: *
Disallow: /recherche/
Sitemap: https://www.exemple.fr/sitemap.xml
Check the CDN layer too
Robots.txt is not the only filter. Cloudflare redefined its default settings on July 1, 2026 (Cloudflare, July 1, 2026). Since September 15, 2026, on new domains, training crawlers and AI agents are blocked by default on pages that display ads. Search crawlers remain allowed. Check the actual setting of your CDN, not just your file.
These trade-offs are covered crawler by crawler in our article on AI crawlers. Note that the llms.txt file neither blocks nor allows any crawler; only robots.txt plays that role.
Which robots.txt for a WordPress site?
WordPress serves a virtual robots.txt as long as no physical file exists at the root. Since version 5.5, it declares its native sitemap there. A robots.txt file uploaded to the root replaces this virtual version. Some SEO plugins also let you edit it from the admin area.
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Disallow: /?s=
Disallow: /search/
Sitemap: https://www.exemple.fr/wp-sitemap.xml
Two rules to follow. Never block /wp-content/ or /wp-includes/ entirely. Stylesheets, scripts and images depend on them, and Google needs them to render the page as a visitor sees it. And adapt the Sitemap line to the plugin you use: the Yoast SEO index, for example, is sitemap_index.xml. See our article on the XML sitemap.
How do you test and monitor your robots.txt?
The old robots.txt Tester in Search Console no longer exists. It has been replaced by the robots.txt report, under Settings, then Crawling. It shows the version fetched by Google, the fetch date and any errors. See our Google Search Console guide.
- Check the response:
curl -I https://www.exemple.fr/robots.txtmust return 200. - Read the robots.txt report in Search Console after each change.
- Inspect a key URL with the URL Inspection tool: it must not be reported as blocked.
- Crawl the site with a tool that respects robots.txt, to list blocked URLs and spot any important page among them.
- Check at every release: a
Disallow: /inherited from staging is the classic mistake of an SEO migration.
Robots.txt is one of the first checks in any technical SEO project. It is short, but a single line can shut down the whole site.
Frequently asked questions
What is robots.txt?
It is a public text file placed at the root of a site. It tells crawlers which parts of the site they may crawl. It follows the Robots Exclusion Protocol, standardized in 2022 as RFC 9309. It manages crawling, not indexing.
How do you find a site’s robots.txt?
Add /robots.txt after the domain name, for example https://www.exemple.fr/robots.txt. A 404 error means the site has no file: crawlers then assume everything is allowed.
What is Googlebot?
Googlebot is Google Search’s crawler. It downloads pages to index them and also feeds AI Overviews and AI Mode. It comes in a smartphone and a desktop version; the smartphone version crawls most sites.
Does robots.txt protect private pages?
No. The file is public, it even points curious visitors to the areas you want to hide, and some crawlers ignore it. A private page is protected with a password or server authentication.
Should you block GPTBot?
It is a business decision. Blocking GPTBot removes your future content from the training of OpenAI’s models, without preventing citation in ChatGPT, which depends on OAI-SearchBot. Decide crawler by crawler, and write the decision down.
Sources
- IETF, RFC 9309, Robots Exclusion Protocol, published in September 2022. Accessed on September 24, 2026.
- Google Search Central, Introduction to robots.txt, updated on December 10, 2025. Accessed on September 24, 2026.
- Google Search Central, How Google interprets the robots.txt specification, updated on August 31, 2026. Accessed on September 24, 2026.
- Google Crawling Infrastructure, Google’s common crawlers (Google-Extended), updated on July 14, 2026. Accessed on September 24, 2026.
- OpenAI, Overview of OpenAI crawlers. Accessed on September 24, 2026.
- Anthropic, Does Anthropic crawl data from the web, and how can site owners block the crawler? Accessed on September 24, 2026.
- Cloudflare, Your site, your rules: new AI traffic options for all customers, published on July 1, 2026. Accessed on September 24, 2026.
Cite this article
, . (2026, September 26). Robots.txt: the practical guide for SEO. Elev8 Lab. https://elev8-lab.fr/en/seo/robots-txt/