Nothing on your site looks broken. The pages open, the menu works, the sitemap validates. And a large share of those addresses has never been fetched by anything, which is why the traffic curve refuses to move.
The sequence is familiar around here. A distributor loads a manufacturer catalog and gains 9,000 part pages overnight. A roofing company covers the suburbs properly and generates 300 combinations of town and service. A freight broker publishes a page per lane. In each case somebody treats going live and being found as one moment. They are two, and the distance between them absorbs most of the wasted effort in industrial and local SEO.
This article works through that distance. Crawl budget and its consumers first, then sitemaps and bulk submission as ways of getting addresses seen, then the limits of the Indexing Hub in the Semalt panel. It ends with arithmetic to run against your own tree.
Published, reached, fetched, kept
Publishing writes a file to a server. Ranking requires four things in order: a crawler learns the address exists, schedules a visit, retrieves the document successfully, then decides the result is worth keeping. Each step fails on its own terms, and only the last produces something a query can match.
The failure modes are dull, which is why they survive for years. An address nothing links to and no sitemap names is effectively invisible. A page retrieved and judged near-identical to 400 siblings is dropped after the fetch. A part reachable only through a faceted filter may be visited once and never again. From a browser all four look the same: the page opens.
Industrial and multi-location sites are unusually exposed because their pages arrive in bulk rather than one at a time. A firm publishing four pages a month notices when one goes missing. A firm that imported a supplier catalog last spring has no idea which 6,000 of its 9,000 part pages were ever retrieved, and nobody has thought to ask.
Crawl budget, and where it goes instead
Crawl budget is the working ceiling on how much of a site a crawler will retrieve in a given stretch of time. Two forces set it: the load your infrastructure carries before response times slip, and how much of the tree is considered worth revisiting. Neither figure is published anywhere, and your own build decisions push both.
Small sites are routinely told this does not concern them. That advice is right about raw URL counts and wrong about composition. A 300-page site squanders its allowance as efficiently as a 300,000-page one, through the same four mechanisms.
- Faceted and parameter addresses. Filter by material, then thread size, then finish, then sort by price, and one catalog page becomes several thousand retrievable addresses. Distribution platforms emit these by default and nobody on the team ever sees them.
- Templated near-duplicates. Thirty suburb pages differing by a town name are each retrieved and each assessed. The cost is charged on all thirty; the value comes back on perhaps four.
- Redirect chains and soft failures. Every hop is a separate retrieval. A discontinued part that answers with a cheerful "product unavailable" page and a success code stays in rotation indefinitely, unlike a clean 404.
- Sluggish responses. Retrieval rate adapts to server behavior. A catalog answering in two seconds gets visited less often than one answering in 300 milliseconds, for identical content.
The practical result is an invisible queue. Whatever is retrieved first gets the attention, and if the opening 500 fetches land on sort-order variants and near-identical town pages, your new service page waits behind them. That is the bottleneck this article is named for. Nothing is penalized and nothing errors; an order of precedence simply got set by accident.
Where the address count comes from around here
Two patterns dominate here, and both are multiplication problems rather than writing problems.
The first belongs to anyone selling a service across the metro. A mechanical contractor knows Naperville and Bridgeport are not one market, so the site gets a page per town and a page per service. Twelve services across twenty-five towns is 300 combinations, plus hubs, from a business with one shop and one fleet.
The second belongs to distribution and manufacturing, where the page counts get genuinely large. A supplier catalog arrives with part numbers and each becomes an address. Add manufacturer landing pages, cross-reference tables between competing numbers and a filter interface across four attributes, and a mid-sized distributor publishes five figures of URLs from a product line one salesperson could describe in a morning.
| What gets built | Addresses | Genuinely distinct content | Realistic outcome |
|---|---|---|---|
| 12 service pages | 12 | High | Retrieved, kept, ranks |
| 25 town hubs | 25 | Medium | Mostly kept |
| Service × town matrix | 300 | Very low | Partly retrieved, largely dropped |
| Catalog by part number | 9,000 | Specification only | Kept where the specification is unique |
| Filter and sort variants | Thousands | None | Pure consumption |
| Freight lanes, origin × destination | 200+ | Low unless priced | A handful survive |
None of this is stupid. Covering the collar counties separately is the right instinct where the competitor set changes between Aurora and Evanston, and a distributor does want the part number findable. The error is generating the whole matrix and hoping quantity stands in for substance. You pay for that in two currencies: retrievals burned on addresses nobody will keep, and attention drawn off the few pages that could have won something.
The question nobody asks before generating
So which pages have earned an address? One test does the work: could this page stand for any other town, or any other part number, with nothing changed but a word? If it could, it has not earned one.
A page about rooftop unit replacement in Schaumburg earns its address by naming the permit process in that village, the building stock it applies to, your shop's response time, and photographs of work done there. If dropping Des Plaines into it produces the neighboring page exactly, nothing local is being targeted. A find-and-replace has been performed, and a crawler reads it that way.
Catalog pages meet the same test on different grounds. A part page earns its keep by carrying the specification, the compatible equipment, stock reality, cross-references to superseded numbers and a price or a route to one. A number, a stock photograph and a request-a-quote button make one of 9,000 identical documents, and it gets treated accordingly.
Addresses worth a retrieval
Pages carrying facts that exist nowhere else on the site, written once and maintained afterwards.
- Four to six towns you genuinely serve
- Your highest-margin services
- Parts with real specifications and stock
Addresses that only consume
Generated output where one variable changes and no verifiable fact appears anywhere on the page.
- The full service × town matrix
- Sort, filter and pagination variants
- Part pages with a number and nothing else
For the mechanical contractor a defensible shape is twelve strong service pages, five town pages for the areas the trucks actually reach, and Spanish versions of the services plus the two towns where that audience lives — around thirty addresses instead of 340. For the distributor it is real product pages for the fast-moving lines, tail parts left inside a filterable listing that is never individually offered for crawling, and the filter interface blocked outright.
A sitemap is a statement, not a formality
Most sitemaps are produced by a plugin and never opened again. Read properly the file does two jobs: it names the addresses you want retrieved, and declares which ones you consider current. On a generated tree it is the cheapest instrument you own, since nowhere else do you state outright what is on offer.
Submitting and parsing sitemap files
For trees where nobody has maintained the address list by hand in years.
- Hand over a file, or name its location. Either the sitemap your domain already publishes, or one assembled for a single job.
- Parsed recursively to three tiers. Where an index names further indexes which name the sitemaps themselves, all three tiers get walked — enough for anything a catalog platform emits.
- One job carries up to 1,000 sitemaps. A whole portfolio goes in as one operation rather than a morning of clicking.
- Two jobs run, twenty wait. Twenty-two sit in flight together and the queue empties in order, unsupervised.
Three tiers sounds like trivia until a catalog is involved. A distributor's index typically splits into one sitemap per manufacturer, each splitting again into blocks of parts — three tiers exactly. Submitting the top file once covers the whole structure, and the same shape turns up in any agency holding a dozen contractor sites.
Direct submission, and what the ceilings really mean
Alongside sitemap jobs, the URL tracker accepts addresses directly. Two ceilings govern it, and how they interact is the thing to work out before a bulk import.
Bulk batches against a daily allowance
A queue with one ceiling per day and another per batch, plus a record of what happened to each address.
- 1,000 addresses a day, counted per account. The allowance belongs to the account and covers every site inside it, which is the figure agencies and multi-brand operators have to plan around.
- A batch may hold 10,000 addresses. Batch size governs how much you hand over at once, not how fast it moves. The queue still empties at the daily rate.
- Delivery over the IndexNow API. GoogleBot and BingBot receive notice that an address changed, rather than finding out on their own timetable.
- Lists handed over in bulk. Stream, the assistant in My SEO, accepts keyword and URL lists in batches, so nobody pastes addresses one at a time.
The arithmetic follows. Ten thousand addresses against a thousand a day is ten days to clear. That is not a defect; it is why ordering matters more than volume. Spend day one on discontinued part numbers and last week's four service pages sit untouched until day nine, when they could have led.
IndexNow is regularly oversold, so state it plainly. It is a notification protocol: participating crawlers learn that an address changed instead of waiting on their own schedule. It compresses the wait before a visit and has no bearing on what follows. Arriving sooner does not make a weak page keepable.
| Instrument | Suited to | Ceiling | What it will not do |
|---|---|---|---|
| Sitemap job | Whole trees and portfolios | 1,000 files, 3 tiers | Order the addresses inside a file |
| Bulk batch | A named list of new or edited pages | 10,000 in a batch | Beat the daily rate |
| Daily allowance | Steady publishing | 1,000 a day, per account | Carry over when unspent |
| IndexNow | Flagging an edit fast | One address at a time | Say anything about the verdict |
What a batch is telling you while it drains
The log runs one line per address rather than per batch, which is what makes it diagnostic. Each line holds the bot visit, its time, a status and the error text where there is one; counters above them track submitted, found and failed. Four patterns account for nearly everything, and each sends you somewhere different.
Requested, never visited
The notice went out and nothing came to look inside the expected window.
- Check robots rules and access control
- Check whether anything links to it
Visited, retrieval failed
Something came and the retrieval broke. That is your infrastructure answering, and the error text says how.
- Timeouts, server errors, redirect loops
- Usually reproducible on demand
Retrieved cleanly, still absent
Every technical step succeeded and the page appears nowhere. This is the honest signal, and the one most often misread.
- A twin of something already stored
- Too slight to warrant a record
Failures piling up
When the failed counter climbs through a batch, one cause is behind it, not fifty.
- Stop the batch and look
- Sending it again repairs nothing
Pattern three costs firms whole quarters precisely because it presents as nothing. No error, no alert, no message: the address was asked for, visited, retrieved without incident and then set aside. Seeing it repeat across a town set or one manufacturer block answers the question about that set, and further submissions will not change the answer. Rewrite two of those pages properly, send those two, and the difference tells you what the other 298 are worth.
Account structure matters here too, since the allowance is shared across the properties inside it. Linked Google account groups, per-site sharing to an outside address and site tags that filter globally let a portfolio be cut by client, by trade or by county, while background workers keep counters current without a manual reload. Fifteen contractor sites at 250 addresses each is under four days of the full allowance — provided somebody chose the order instead of sending it all at once.
Questions that come up during the first batch
We submitted 400 addresses and 120 got indexed. Did something fail?
No. If the log shows visits and clean retrievals, the submission did exactly its job. The other 280 pages received a verdict, and on a generated town or catalog tree that proportion is ordinary. It is information about those pages, not about the queue.
Should the whole catalog go in, or only what changed?
Only what is new or genuinely edited. Re-sending untouched parts spends the day's allowance for nothing, and what goes unspent does not roll forward. Run a standing queue of real additions rather than a monthly push of everything.
Our filter interface generates thousands of addresses. What do we do with them?
Keep them out of the sitemap first, since that list is the one you control directly, then block the parameter patterns at the robots level and stop linking filtered views as ordinary navigation. The aim is not tidiness: retrievals spent on sort orders are retrievals not spent on the parts you sell.
Do the Spanish pages need their own sitemap?
One file can carry both, and at this size that is easier to maintain. What matters much more is that the pages were written rather than passed through a translator, and that they cover the areas where that audience actually lives. A mirrored tree doubles the address count and usually gets neither half kept.
How long before we conclude a page will never be indexed?
Let the batch drain, then wait another fortnight. If the log shows a clean retrieval and the page is still missing after a month, more requests change nothing. The page is the variable.
The calculation worth running first
Take your own figures. Count the services you actually sell. Count the towns you want a page for. Multiply those, then add the catalog: parts, manufacturers, cross-references. Now count, honestly, how many of those pages somebody on your team could write one specific paragraph about — a permit rule, a lead time, a compatibility note, a price.
The second number is the real one, and the first is usually five to fifty times larger. Everything in the gap is retrieval you will get nothing back for, queued at a thousand addresses a day and judged by a standard that never counts how many pages you published.
Discovery comes before that timeline, not alongside it. Four to eight weeks to first measurable movement presumes the pages already sit in the index; where half a tree was never discovered, the campaign is not slow but incomplete. Settling discovery first is what gives the remaining work — keyword selection and link placement across a partner network exceeding 230,000 sites — something to act on. The same reduction shapes the services we run, and further walkthroughs sit on our blog.
The firms that get this right around here are not the ones with the largest page counts. They are the ones that picked four or five towns they can genuinely claim, gave fast-moving parts real specifications, kept the filter interface out of the crawl, and checked within a week of launch that every one of those pages had actually been retrieved.
To learn which of your addresses have ever been visited, connect the domain and point a sitemap job at the tree as it stands: open the Semalt dashboard and start an indexing job. The first finding worth having is seldom a technical fault. It is the block of pages that came back clean months ago and has been sitting outside the index ever since.