Analytics, SEO and growth · Format

robots.txt

A small text file at the root of a site that tells search engine and AI crawlers which parts they may visit and which to leave alone.

Telling crawlers · updated

How it works

It lives at /robots.txt and groups rules by crawler: a User-agent line (Googlebot, Bingbot, GPTBot, or * for everyone) followed by Allow and Disallow lines with paths, plus an optional Sitemap line. The format was standardised as RFC 9309 in 2022. Reputable crawlers obey it, but it is a request, not a lock: it offers no security, and anyone can read the file.

It controls crawling, not indexing. A blocked page can still appear in results, without a description, if other sites link to it; to keep a page out of search, let it be crawled and add a noindex robots meta tag or X-Robots-Tag header. AI companies often use separate crawlers for training and for search (OpenAI has GPTBot and OAI-SearchBot), so each can be allowed or blocked on its own.

robots.txt vs the alternatives

More in Analytics, SEO and growth

Telling crawlers

All 20 Analytics, SEO and growth terms

Crafted in the dark. Shipped to the world.

Tell us what you are building. You get a private project space with a proposal and a line-by-line quote within a day.