Crawl Budget: The Resource Google Just Confirmed Is Shared
The Budget No One Told You Was Shared
On July 22, 2026, Google quietly rewrote its crawl budget documentation. No blog post, no announcement, no ranking update to react to. Just a rewrite of a page most site owners will never open, buried under “Optimize crawling performance” in the Search Central docs.
It’s worth reading anyway. Not because it changes how crawling works – Google is explicit that it doesn’t – but because it writes down, precisely, something that was previously left to inference. And one line in particular changes the conversation around AI crawlers in a way that has nothing to do with rankings and everything to do with who gets to be read at all.
What crawl budget actually is
Crawl budget is not a mysterious algorithmic score. It’s closer to an operational limit: the set of URLs that Google can crawl on your site without straining your server, combined with the set of URLs Google actually wants to crawl based on how much your content changes and how much it matters.
Google calls the first part the crawl capacity limit – essentially how many connections your server can hold open for Google without degrading. The second part is crawl demand – how often Google feels it’s worth coming back, shaped by freshness, popularity, and how much of your content looks unique rather than duplicated.
Multiply the two together and you get a site’s crawl budget: not a number you can check anywhere, but a behaviour you can observe in how quickly new or updated pages get picked up.
Most sites never need to think about this. If your pages tend to get crawled the same day they’re published, this entire topic is irrelevant to you – that’s Google’s own framing, and it’s honest. This matters for large catalogues, frequently updated sites, and any domain where Search Console shows a large share of URLs stuck in “Discovered – currently not indexed.”
What actually changed
Three things in the July rewrite are worth sitting with.
Every site starts at the same conservative baseline. Google now states explicitly that new and established sites begin with the same default crawl capacity limit. The limit rises only if there’s demand to crawl more and the site stays healthy – meaning consistent response times and no server errors. There’s no head start for age or size. There’s a starting line, and then there’s how well your server holds up under repeated visits.
Crawl health is now framed around consistency, not just speed. A site that responds reliably, with stable or improving latency and Time-to-First-Byte, earns a higher capacity limit over time. A site that slows down under load, throws 5xx errors, or returns rate-limiting signals gets crawled less. This is the same argument that infrastructure quality has always made for itself – it’s just now written into the official mechanism rather than left as an inference from case studies.
Crawl capacity is shared across every one of Google’s crawlers. This is the line that matters most. Googlebot, AdsBot, Google-Shopping’s crawler, and Google-Extended – the crawler that gathers training and grounding data for Gemini and AI Overviews – all draw from the same finite capacity pool on your site. High demand from one can reduce what’s available for the others.
That single sentence quietly resolves a debate that’s been running across technical SEO circles for the past year: whether AI crawlers compete with search crawlers for the same resources, or operate independently. Google has now answered it. They compete.
Why this matters more than it looks like it should
If you’ve read anything here about how search is shifting from ranking pages to grounding answers, you’ll recognise where this is going.
A system that generates answers instead of showing a list of links depends on being able to revisit sources often enough to know they’re still accurate. Freshness isn’t a ranking factor in that world – it’s a precondition for being citable at all. And freshness, mechanically, is downstream of crawl demand and crawl capacity: Google can’t reflect an update it hasn’t fetched.
Now that Google confirms its AI-oriented crawlers pull from the same capacity pool as its search crawlers, an inefficient site isn’t just crawled less for ranking purposes. It’s crawled less, period – across every Google product that depends on visiting your pages, including the ones deciding whether to cite you in a generated answer.
A slow, poorly cached, duplicate-heavy site doesn’t just rank worse. It becomes structurally harder to keep current in every system that Google runs on top of its crawl infrastructure. That’s a much larger cost than a ranking penalty, and it’s invisible until you go looking for it.
What actually earns more crawl budget
Google’s own guidance narrows down to a short list of things that genuinely move the needle, and none of them are exotic.
Serve pages fast and consistently
Response time and stability are now explicitly tied to the capacity limit. Faster, more consistent servers get crawled more. This is precisely the argument for infrastructure quality that shows up whenever performance and SEO intersect – it just now has an official mechanism behind it rather than a correlational case study.
Support conditional requests
Returning a 304 status when a page hasn’t changed since Google’s last visit tells the crawler to reuse its cached copy instead of re-downloading the page. It’s a small technical detail that directly reduces the cost of every crawl, which frees up capacity for the pages that did change.
Keep your URL inventory honest
Duplicate content, parameter variants, and pages you don’t actually want indexed all consume crawl budget without returning any value. Consolidating duplicates and using robots.txt to block what should never be crawled – not what you’re trying to temporarily deprioritise – keeps the crawler focused on what matters.
Don’t rely on noindex to save budget
This one catches people out. A noindex tag doesn’t stop Google from requesting the page – Google still has to fetch it to see the tag before dropping it. If the goal is to stop the crawl entirely, robots.txt is the correct tool. Noindex only removes the page from the index; it does nothing to protect crawl budget.
Fix soft 404s and avoid long redirect chains
Both waste repeated crawl cycles on pages that will never return anything useful. A hard 404 or 410 is a clear, one-time signal. A soft 404 – a page that returns 200 but functionally has no content – keeps getting revisited indefinitely.
An informal term Google eventually had to define
Crawl budget followed the same trajectory as another concept discussed on this site before: it started as an explanatory heuristic inside the SEO community, used informally for years before Google ever published an official definition. There’s no Wikidata entity for it, no formal academic origin. It exists because practitioners needed a word for a pattern they kept observing, and Google eventually wrote its own definition once the term had already taken hold.
That’s a genuine grounding gap – a widely used concept with no canonical, machine-readable entity behind it. It’s exactly the kind of gap the DSH concept graph exists to close: a verified entity descriptor, sourced directly from Google’s own infrastructure documentation rather than secondary commentary, with its relationship to technical SEO, site auditing, and E-E-A-T made explicit rather than left implicit.
The quiet argument underneath all of this
None of this is really about crawl budget. It’s about what happens to a site that treats its own infrastructure as an afterthought.
A site that loads fast, avoids duplication, handles conditional requests correctly, and stays healthy under repeated visits earns more attention from every crawler Google runs – including the ones now confirmed to feed the systems generating answers instead of listing links. A site that doesn’t gets less of everything: less frequent reindexing, slower recognition of updates, and a smaller share of a capacity pool that Google has just confirmed is shared, finite, and increasingly contested.
The documentation didn’t change how crawling works. It just stopped leaving that part to inference.
Título propuesto: Crawl Budget: The Resource Google Just Confirmed Is Shared
Meta description: On July 22, 2026, Google rewrote its crawl budget documentation and confirmed something practitioners suspected: AI crawlers and search crawlers share the same finite capacity pool. What that means for site infrastructure.
Abstract: Analysis of Google’s July 2026 rewrite of its crawl budget documentation, focused on the confirmation that crawl capacity is shared across all of Google’s crawlers – including those feeding AI Overviews and Gemini grounding. Connects crawl infrastructure quality to freshness and citability in generative search, and traces crawl budget’s origin as an informal SEO-community term later formalised by Google, mirroring the same grounding gap addressed by the DSH entity descriptor for the concept.


