Robots.txt generator

Robots.txt generator

Configure which crawlers can access which pages of your site and generate a ready-to-deploy robots.txt file.

Crawl rules

What is robots.txt and how does it work?

The robots.txt file is a plain-text file placed at the root of your website (e.g. https://example.com/robots.txt) that instructs web crawlers (search engine bots, AI scrapers, and other automated agents) about which pages they should or should not access. It is governed by the Robots Exclusion Protocol (REP), a 1994 standard that has been adopted by virtually all major search engines.

Robots.txt works on the honour system, it is a polite request, not a technical barrier. Well-behaved crawlers like Googlebot, Bingbot, and most legitimate bots respect these rules. Malicious scrapers ignore them. Therefore, robots.txt should never be used to hide sensitive information, use authentication and server-side access controls for that.

Common uses for robots.txt include: blocking the /admin/ path from being indexed, preventing search engines from crawling session-based or tracking URLs (like ?ref= parameters), keeping staging or development paths off the index, and controlling AI training crawlers (GPTBot, CCBot, anthropic-ai) that have proliferated since 2023. The Crawl-delay directive asks bots to wait a specified number of seconds between requests, useful for high-traffic sites where aggressive crawling causes server load.

Important limitations

  • A Disallow directive prevents crawling but does not de-index already-indexed pages.
  • To remove already-indexed pages, use noindex meta tags or Google Search Console's URL removal tool.
  • Robots.txt is publicly visible, never put sensitive paths there you do not want known.

Frequently asked questions

Where do I put my robots.txt file?

At your website's root: https://yourdomain.com/robots.txt. It must be served as plain text (Content-Type: text/plain). Upload it to your public_html or www root directory.

Does robots.txt affect SEO?

Yes, blocking important pages from crawling prevents them from ranking. Accidentally blocking /wp-content/ or CSS/JS files can also hurt Core Web Vitals scores because Google can't assess page experience.

How do I block AI training bots?

Add rules for GPTBot, CCBot, anthropic-ai, and Google-Extended. Note that not all AI crawlers identify themselves honestly or respect robots.txt.