robots.txt is a text file at the root of your domain that communicates crawl instructions to web crawlers. It’s the fastest way to tell Google what not to crawl — but it’s a blunt instrument that requires care. Blocking the wrong paths hides content from Google; failing to block the right paths wastes crawl budget on content you don’t want indexed.
A basic understanding of robots.txt syntax is insufficient for most sites. Understanding per-crawler directives, crawl rate management, the relationship between disallow and noindex, and common traps makes the difference between a robots.txt that helps and one that silently causes problems.
The Robots.txt Protocol
The Robots Exclusion Protocol defines how robots.txt works:
- Located at
https://example.com/robots.txt(always at root domain, not subdirectory) User-agent:specifies which crawler the following rules apply toDisallow:specifies paths the crawler should not accessAllow:specifies paths within a disallowed directory that should still be accessible*wildcard matches any sequence of characters
Example:
User-agent: Googlebot
Disallow: /admin/
Disallow: /checkout/
Allow: /admin/public-report/
Crawl-delay: 10
User-agent: *
Disallow: /private/
Per-Crawler Directives
A crucial capability: different rules for different crawlers. Common use cases:
Blocking specific crawlers entirely:
User-agent: AhrefsBot
Disallow: /
This blocks Ahrefs’ crawl bot completely while allowing Googlebot.
Allowing Googlebot but not other bots:
User-agent: Googlebot
Allow: /
User-agent: *
Disallow: /
Blocks all crawlers except Googlebot. Useful for sites that only care about Google indexation and want to prevent other scrapers.
Differentiating Googlebot from Google image crawlers:
User-agent: Googlebot
Disallow: /internal-docs/
User-agent: Googlebot-Image
Disallow: /
Prevents Google from indexing your images (Googlebot-Image handles image search indexation) while allowing regular page crawling.
What Disallow Does (and Doesn’t Do)
Critical misunderstanding: Disallow prevents crawling, not indexation.
If an external site links to a page that’s in a disallowed path in robots.txt, Google may still index the page — it just can’t crawl it. Googlebot won’t see the page’s content, but it may create a “URL known, not crawled” entry in its index, which can appear in search results as a result with no snippet (“A description for this result is not available”).
For pages you want to never appear in search results, use noindex in the page’s <meta> tags — not just robots.txt disallow.
Combination pattern for complete exclusion:
- Allow crawling (don’t disallow) but add
<meta name="robots" content="noindex"> - Or: disallow AND add the URL to GSC’s URL removal tool
Crawl Rate Management
The Crawl-delay directive is supported by some crawlers (including Bing/Bingbot, not Google):
User-agent: Bingbot
Crawl-delay: 10
This tells Bingbot to wait 10 seconds between requests, preventing aggressive crawling from impacting server performance.
For Google, crawl rate management is handled in GSC (Settings → Crawl stats → Change crawl rate). Google doesn’t honor Crawl-delay in robots.txt for Googlebot.
Sitemap Declaration
robots.txt can declare the location of your XML sitemap(s):
Sitemap: https://example.com/sitemap.xml
Sitemap: https://example.com/news-sitemap.xml
This isn’t a crawl directive — it’s a discovery shortcut. Crawlers that process robots.txt will find your sitemap even if they haven’t been directly submitted it via GSC or search engine webmaster tools. Best practice: always include Sitemap declarations in robots.txt.
Common robots.txt Mistakes
Disallowing CSS and JavaScript: Some sites block /assets/, /js/, or similar paths to reduce crawl load. Googlebot needs CSS and JavaScript to render pages for indexation. Blocking these creates rendering failures — pages appear unstyled to Google, affecting quality assessment. Never disallow CSS/JS assets unless they contain private content Google shouldn’t see.
Disallowing the entire site:
User-agent: *
Disallow: /
This blocks all crawlers from everything. Usually added by a CMS during staging setup and forgotten when the site goes live. A site with this robots.txt won’t be crawled or indexed.
Blocking parameter URLs without considering their use: Blocking /*? prevents crawling of all parameterized URLs, including tracking parameters (good) but also potentially navigation states, search results (which may be fine to block), and any other parameter the site uses.
Not testing after changes: robots.txt syntax errors or unintended path patterns can silently block important content. Use Google’s robots.txt Tester in GSC to verify rules work as intended before and after changes.
The Allow Directive for Subdirectory Exceptions
When you disallow a parent directory but need specific subdirectories to remain crawlable:
User-agent: Googlebot
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
This blocks all of /wp-admin/ except admin-ajax.php (which powers front-end AJAX functionality). The specific Allow overrides the broader Disallow for that path.
Precedence rule: longer, more specific paths take precedence over shorter, broader paths. /wp-admin/admin-ajax.php is more specific than /wp-admin/ — the Allow wins.