Crawl budget is not a concern for sites with under 1,000 pages where Google can crawl everything in a single visit. It becomes critical for e-commerce sites with millions of product pages, news sites with high publication volume, or any site where faceted navigation, URL parameters, or dynamic content generation creates an enormous crawlable URL space.
Understanding Crawl Budget
Google’s crawl budget has two components:
Crawl rate limit: How fast Googlebot crawls your site without overwhelming your server. Google automatically adjusts this based on server response time — slow responses cause Google to slow down. You can reduce (but not increase) this limit in Google Search Console.
Crawl demand: Google’s desire to crawl your URLs based on their popularity and freshness. High-traffic, frequently updated pages get crawled more often. Low-traffic, rarely updated pages get crawled infrequently.
The practical crawl budget is the product of these: Google will crawl a certain number of your pages per day, prioritized by demand signals. For most sites, this is sufficient. For large sites, you may have more crawlable URLs than Google crawls in any given period — new content may sit for days or weeks before being indexed.
Identifying Crawl Budget Issues
Signs of crawl budget problems:
- New content takes weeks to appear in search results
- Google Search Console shows “Discovered but not crawled” for many important pages
- Large number of 200-status pages with no organic traffic (never indexed despite being crawlable)
- High ratio of low-quality pages to high-quality pages
Diagnosis tools:
- Google Search Console → Pages → “Discovered, currently not indexed” — these pages Google knows about but hasn’t gotten to
- Server log analysis (Screaming Frog Log File Analyser, ELK stack) — shows actual Googlebot crawl frequency per URL
- Screaming Frog crawl audit — maps all crawlable URLs including those you may not have intended to create
Common Crawl Budget Wasters
Faceted navigation parameters: E-commerce sites with filter navigation generate enormous URL spaces:
/products/shoes?color=red&size=10&brand=nike/products/shoes?size=10&brand=nike&color=red(same page, different parameter order)/products/shoes?sort=price-asc
These parameter combinations create millions of near-duplicate URLs. Solutions:
robots.txtdisallow for parameter-based pages that shouldn’t be crawledrel="canonical"on all parameter URLs pointing to the base collection page- Google Search Console’s URL Parameters tool (partially deprecated but still functional) to mark parameters as crawl-variant or non-variant
Session IDs in URLs: example.com/page?sessionid=abc123 — session IDs create a unique URL for each user session. These should be handled server-side, not in URLs. Block via robots.txt or canonical to the sessionless URL.
Pagination beyond page N: For e-commerce with 1,000+ products, deep pagination pages (page 50, page 100) have minimal value. Crawl and index the first 3-5 paginated pages; block deeper pagination.
Infinite scroll: JavaScript-based infinite scroll that appends content without URL changes is invisible to crawlers — Googlebot sees only the first screen’s content. If you use infinite scroll, implement URL updates as users scroll (History API) or provide a paginated fallback.
Staging or development content leaked to production: QA pages, A/B test variants, and development endpoints accidentally indexed.
Robots.txt for Crawl Budget Management
robots.txt disallows crawling — it does not de-index pages. Pages blocked by robots.txt but indexed can remain in the index. For true exclusion from index, use noindex meta tags or HTTP headers.
Effective robots.txt patterns for crawl budget:
User-agent: Googlebot
# Block faceted navigation parameters
Disallow: /*?*color=
Disallow: /*?*size=
Disallow: /*?*sort=
# Block cart, account, search result pages
Disallow: /cart
Disallow: /account/
Disallow: /search
# Block staging paths that may have leaked
Disallow: /staging/
Disallow: /test/
Caution: robots.txt mistakes can disallow critical pages or entire site sections. Test changes in Google’s robots.txt Testing Tool before deploying.
XML Sitemap as Crawl Guidance
XML sitemaps don’t guarantee crawling, but they signal to Google what you consider important and provide discovery for URLs not well-linked internally.
Sitemap best practices for large sites:
- Segment sitemaps by content type:
sitemap-products.xml,sitemap-blog.xml,sitemap-categories.xml - Only include URLs that should be indexed: don’t include noindexed pages, parameter URLs, or low-value pages
- Use
<lastmod>for pages that update regularly — accurate lastmod signals help Google prioritize crawl of fresh content - Maximum 50,000 URLs per sitemap file; use sitemap index files to reference multiple sitemaps
Sitemap and index coverage correlation: Check Search Console’s Sitemaps report against the total indexed count. If you submit 10,000 URLs via sitemap but only 5,000 are indexed, the 5,000 non-indexed URLs need investigation — are they thin? Blocked? Failing crawl?
Internal Linking for Crawl Efficiency
Google discovers pages primarily through internal links. Internal linking architecture directly controls how crawl budget is distributed:
Deep pages: Important pages buried many clicks from homepage receive less crawl frequency. Lift critical pages by linking to them from the homepage or high-traffic pages.
Orphan pages: Pages with no internal links are invisible to crawl. Ensure every important page has at least one internal link from a crawled page.
Link equity distribution: PageRank-equivalent signals flow through internal links. Pages receiving many internal links get crawled more frequently (because they appear more important) and rank better.
Crawl-efficient site architecture:
- Flat structure preferred: most pages reachable in 3 clicks
- Pagination for large collections: page 2 links to page 3 (enabling sequential discovery)
- Category → subcategory → product link chains: each level links to the next