FR, version française
Home / Blog / SEO / Crawl budget: understanding and optimising crawling

Crawl budget: understanding and optimising crawling

Crawl budget is one of the most quoted notions in technical SEO, and one of the most misunderstood. Google defines it precisely: a capacity limit, set by the health of your server, and a crawl demand, set by the interest of your pages. Its documentation, updated in July 2026, also says who should care: very large sites and those that change every day. This article helps you find out whether your site is concerned, read crawl stats and logs, then apply the levers that work. It also rules out the reflexes that free up no budget, such as noindex.

Key takeaways

Crawl budget is the time and resources Googlebot devotes to crawling a site. It combines two components. The capacity limit depends on the health of the server. Crawl demand depends on the size, freshness and quality of the pages. It only becomes a real issue on large sites.

  • Google targets sites with 1 million pages or more, or 10,000 pages or more updated every day: orders of magnitude, not thresholds.
  • Noindex frees up no budget: the page is crawled anyway.
  • The levers that count: fewer useless URLs, a fast server, an up-to-date sitemap.

What is crawl budget?

The web is too vast for Google to crawl every public URL. So it shares out its crawling time between sites. This is the crawl budget (Google, Crawl Budget Management, July 22, 2026). Google describes it as the meeting point of two components.

ComponentWhat it measuresWhat makes it vary
Crawl capacity limit (or hostload)The time your server spends keeping connections open for Google: number of parallel connections and durationFast, stable responses raise it; slowdowns, 5xx errors or 429s lower it
Crawl demandGooglebot’s interest in your URLsSite size, update frequency, quality and relevance of the pages, perceived URL inventory
Source: Google documentation on crawl budget management, accessed on September 24, 2026.

The July 2026 update clarifies a useful point. All sites start with the same capacity limit, deliberately cautious. Google then raises it if it has reasons to crawl more and if the site stays healthy.

Warning

Crawled does not mean indexed. After crawling, each page is evaluated, consolidated with its duplicates, then kept in the index or not. An indexing problem should first be treated as such: see our article on Google indexing.

Does crawl budget concern your site?

For the vast majority of sites, no. Google reserves its guide for three profiles:

  • large sites, in the order of 1 million unique pages or more, whose content changes about once a week;
  • medium or large sites, in the order of 10,000 unique pages or more, whose content changes every day;
  • sites with a large share of URLs listed as “Discovered – currently not indexed” in Search Console.

Google adds that these numbers are a rough estimate, “not exact thresholds” (same documentation). It also gives a simple test. If your pages are crawled on the day they are published, this guide does not concern you. An up-to-date sitemap and a regular look at the indexing report are enough.

Remember

A 200-page brochure site almost never has a crawl budget problem. Nor does a blog with 1,000 articles. If its pages are not indexed, the cause lies elsewhere: perceived quality, duplicates, orphan pages, a technical block.

How can you tell if Googlebot is wasting your crawl budget?

The Crawl Stats report

In Search Console, the Settings menu, then Crawl stats, shows 90 days of Googlebot activity. The report only exists for root-level properties, and Google considers it of little use below 1,000 pages. Four readings are worth making:

  • by response: a high share of 301s, 404s or 5xx signals wasted time;
  • by file type: too many requests on scripts or images at the expense of HTML;
  • by purpose: the share of discovery (new URLs) compared with refresh;
  • host status: robots.txt availability, DNS, server connectivity.

Server logs

Logs record every request received by the server. They are the only source that tells you what Googlebot really does, URL by URL. Take two precautions. First, check that the visits attributed to Googlebot really come from Google. A reverse DNS lookup or the published IP ranges confirm it (Google, verifying crawlers, March 20, 2026). Then cross-check the crawled URLs with those in your own crawl. A strategic page rarely visited, or thousands of heavily visited parameter URLs, show where the budget goes.

For tools, Screaming Frog’s Log File Analyser is free up to 1,000 log lines and one project. Beyond that, the licence costs £99 a year (price checked on September 24, 2026). The full method is explained in our guide to log file analysis.

Which levers really improve crawling?

Google’s documentation lists specific practices. Here they are, from the most structural to the most technical:

  1. Reduce useless URLs: merge duplicates (see duplicate content), limit combinations of filters and sort orders.
  2. Block what is useless, meaning what should never be crawled: sorted versions of the same list, infinite scrolling that repeats linked pages. The right tool is the robots.txt file, for a lasting block.
  3. Return 404 or 410 for deleted pages: Google then gradually reduces how often it crawls them (Google, HTTP status codes and network errors, updated on February 4, 2026). Soft 404s, errors served with a 200, keep consuming budget.
  4. Shorten redirects: each hop is one more request. Internal links should point to the final URL.
  5. Keep the sitemap up to date, with a reliable lastmod tag. Google only uses this date if it is consistently accurate (Google, build a sitemap, updated on July 8, 2026). See our article on the XML sitemap.
  6. Speed up the server and manage HTTP caching: a stable response time raises the capacity limit. The 304 (Not Modified) code lets Google reuse its copy without downloading the page again.

Finally, crawling follows links. A page five clicks from the home page is usually discovered later than a page linked from a section. Internal linking remains the simplest way to show Googlebot what matters. The rule “three clicks at most for a strategic page” is an agency convention. Google does not state it.

To get more budget, Google cites only two routes. The first: add server resources if URL inspection reports “Hostload exceeded”. The second: improve content quality, meaning its popularity, its value to the user and its uniqueness.

What does not free up crawl budget

Several common reflexes have no effect, or the opposite effect. Google’s documentation explicitly rules them out.

MisconceptionWhat Google says
“I set weak pages to noindex to save budget”False. Google still requests the page, then drops it when it sees the noindex: the crawling time is spent. Noindex is used to block indexing, not crawling.
“I temporarily block one section so that Google crawls the other”False. Google does not reallocate the freed budget, unless your site is already at its capacity limit. robots.txt is for blocking, over the long term, what should not be crawled.
“The more Google crawls, the better I rank”Not documented. No Google page presents crawl frequency as a ranking signal. More frequent crawling mainly means your changes are taken into account sooner.
“I return errors to calm Googlebot down during an overload”Possible, but only for a few hours to one or two days, with 500, 503 or 429 codes. Beyond that, Google may remove pages from the index (Google, reduce the crawl rate, updated on December 18, 2025).

Crawl budget is a chapter of technical SEO, not a goal in itself. On a large site, it is managed with logs and crawl stats. On the others, keeping an eye on the indexing report is enough.

Frequently asked questions

What is crawling?

Crawling is a bot’s visit to the pages of a site. It downloads their content and discovers new links. It is the first step before indexing and then ranking in a search engine.

What is a crawler?

A crawler is a program that travels the web from link to link. Googlebot works for Google, Bingbot for Bing, and AI answer engines have their own too. SEO tools such as Screaming Frog use the same principle to audit a site.

What is the difference between crawling and scraping?

Crawling explores pages to discover and index them. Scraping extracts specific data from a page (prices, texts, contact details) to reuse it. A search engine crawls; a scraper can ignore robots.txt, which search engines respect.

How can you increase your crawl budget?

Google cites two routes: sufficient server resources if the site is saturated, and better-quality content. Reducing useless URLs does not raise the budget, but it concentrates it on the pages that matter.

Sources

  1. Google Crawling Infrastructure, Crawl Budget Management For Large Sites, updated on July 22, 2026. Accessed on September 24, 2026.
  2. Google Search Central, How HTTP status codes affect Google’s crawlers, updated on February 4, 2026. Accessed on September 24, 2026.
  3. Google Search Central, Reduce the Google crawl rate, updated on December 18, 2025. Accessed on September 24, 2026.
  4. Google Crawling Infrastructure, Verify Google crawlers and fetchers, updated on March 20, 2026. Accessed on September 24, 2026.
  5. Google Search Central, Block Search indexing with noindex, updated on December 10, 2025. Accessed on September 24, 2026.
  6. Google Search Central, Build and submit a sitemap, updated on July 8, 2026. Accessed on September 24, 2026.
  7. Screaming Frog, Log File Analyser, product and pricing page. Price checked on September 24, 2026.
claude-editeur Avatar

Digital marketing, SEO and GEO consultant

More about the author

Article checked and updated by the author. Sources consulted on the date shown.

Cite this article

, . (2026, September 26). Crawl budget: understanding and optimising crawling. Elev8 Lab. https://elev8-lab.fr/en/seo/crawl-budget/