What robots.txt does — and what it does not
The robots.txt file, placed at the root of your site (www.example.com/robots.txt), tells crawlers which parts of the site they may explore. It is useful to keep them away from pages with no value in search results: shopping cart, customer account, internal search results, back office. It follows a standard syntax (RFC 9309) that the main search engines respect.
Be careful: robots.txt prevents crawling, not indexing. A blocked page can still appear in Google if other sites link to it. To keep a page out of the results, let it be crawled and add a noindex tag, or protect it with a password.
Common mistakes
- Blocking the whole site by accident: a “Disallow: /” left over from development is enough to make a site disappear.
- Blocking CSS and JavaScript: Google needs them to render your pages correctly.
- Forgetting the sitemap: the Sitemap line helps search engines discover all your pages.
The AI crawler options let you refuse the use of your content by some AI services; each company documents the name of its crawler, and the list changes regularly. After publishing the file, check that your important pages can still be crawled with the SEO analyzer, and that old addresses redirect properly with the redirect checker.
Are you an agency? We work white-label →
Frequently asked questions
Does robots.txt stop a page from appearing in Google?
No. It stops crawling, not indexing. To keep a page out of the results, use a noindex tag or password protection.
Where do I put the file?
At the root of the domain: https://www.example.com/robots.txt. One file per domain and subdomain.
Can I block AI crawlers?
Yes, for the crawlers that respect robots.txt, by naming them (GPTBot, ClaudeBot, Google-Extended…). The list changes regularly.
