A page outside the index does not lose to the competition — it is not even in the race. This article shows why discovery, crawling and indexing are three separate phases, what the Indexing Hub in the Semalt panel makes measurable about them, and which limits that instrument deliberately sets.
Almost every SEO plan starts with copy, keywords and links, understandably: that is where the work is visible and discussable. The step before it goes undiscussed. Does the search engine know your page at all, and has it decided to keep it? While that stays open, you are tinkering with something whose existence the search engine does not acknowledge.
Nor is this a matter for large portals alone. An Amsterdam events platform with thousands of performances, a creative-sector webshop with a broad range, a software company with bilingual documentation: in each case addresses appear at a rate nobody tracks by hand. That a backlog exists usually shows only when organic traffic stalls while the editors publish every week.
Three phases often mistaken for one
The most persistent misconception treats indexing as one event that either happens or does not. In reality there are three consecutive phases, each with its own logic and its own kind of blockage. A problem arising in the first phase cannot be repaired in the third, however many submissions you throw at it.
- Discovery. The search engine learns that an address exists, through internal links, a sitemap, an external link or an active notification. Without a route to it, the page does not exist in its world.
- Crawling. The bot actually fetches the address. How many fetches your domain gets in a period is called crawl budget and follows from server performance, earlier response quality and the estimated value of the domain.
- Indexing. The fetched page is assessed and kept, or not. This decision is about content, duplication and signals that either agree or contradict.
Weeks can pass between those three acts. An address can be known for months without ever being fetched; it can be fetched and still refused. Hold on to that distinction and you ask the only useful question: at which step is it stuck? A broader overview of the modules sits in the description of the rebuilt Semalt panel; here it is purely the indexing part.
| What you observe | Phase where it goes wrong | What you do |
|---|---|---|
| No bot has ever visited the URL | Discovery | Create an internal link, extend the sitemap, notify actively |
| Visit logged, but with a server error | Crawling | Check infrastructure and response times under load |
| Visit succeeded, page stays out of the index | Indexing | Bring content, canonical URL and internal signals into line |
| Only deep pages lag behind | Discovery and crawling | Reduce click depth, clean up pagination and filters |
| New pages only appear after weeks | Discovery | Use active notification instead of waiting for a routine visit |
Six causes that keep repeating
In audits the same causes return with a predictability that is almost comic. The six below cover the great majority of cases, ordered by the phase in which they arise.
| Cause | Phase | How you recognise it | What fixes it |
|---|---|---|---|
| Orphan page without inbound links | Discovery | Answers with status 200, appears in no sitemap or navigation | Place it in a relevant section and in the sitemap |
| Excessive click depth | Discovery | Twelve clicks from the homepage, reachable only through pagination | Flatter structure, hub pages, better internal linking |
| Duplicate content through parameters | Indexing | The same article via three paths, with sort and filter variants | Pick one canonical URL and use it consistently |
| Contradictory canonicals | Indexing | Pagination points at page one, targets redirect or are missing | Point the canonical at an existing, indexable URL |
| robots.txt confused with noindex | Crawling | Blocked URL stays in the index, because the noindex is unreadable | Open access and let the noindex be read |
| Slow or unstable delivery | Crawling | Response times climb under load, 5xx errors during peaks | Sort out caching, hosting and error handling |
Two of these six deserve a separate warning. The confusion between robots.txt and noindex costs weeks at every launch: a blocked address is not read, so neither is the instruction inside it. And a canonical is a recommendation, not an order; the search engine may pick a different variant than you intended.
Amsterdam adds a complication of its own. A company that publishes its offering in Dutch and in English — the rule here rather than the exception — doubles its stock of addresses. The crawl budget does not double with it. Anyone managing large volumes without control over response times therefore starts not at submission but at the technical foundations.
The Indexing Hub: the limits are the design
Three functions sit together in the Indexing Hub: submitting addresses, reading sitemaps and recording what the bots then do with them. The most useful part is not the capabilities but the limits, because they determine what is achievable within which period.
The smallest number is the most interesting. A thousand addresses a day is generous for a company site of seventy pages, but with a catalogue of tens of thousands it becomes a scarce resource to be allocated, much like a media budget. That is where the gain of the limit sits: it forces a choice that would otherwise be postponed indefinitely.
Bulk submission of addresses
For moments when many URLs change at once: a migration, a redesign, a new season.
- One send, no portions. The whole list goes out in one go; splitting it into manageable chunks falls away.
- The daily budget stays the brake. Handing over 10,000 addresses at once still means ten days of processing. Batch size therefore does not change the lead time.
- The ordering is the real decision. With a tight budget, the ranking is the only thing you can still steer, and it determines the return of week one.
- For exceptions, not for the routine. The daily publishing pace of a normal site never comes near this ceiling.
Sitemap processing
For collections already organised in sitemap indexes, for example by language or by category.
- Upload or reference. Both an uploaded file and a public address are accepted, so you can also run a set-up through before it goes live.
- Chains up to three layers deep. If an index points at a further index, that chain is expanded to three layers. Nest deeper than that and you are better off flattening the set-up.
- A thousand files in one job. That covers the complete split of a large site by language, section or content type in a single round.
- Two at a time, twenty waiting. Enough for an agency with a handful of Amsterdam clients, provided somebody guards the order: Friday's launch belongs ahead of routine maintenance.
Both entry points end in the same observation and share the same daily budget. How the indexing module from Semalt relates to the analytics and campaign parts is set out in the platform overview.
Pushing instead of waiting
Traditionally the search engine fetches when it suits the search engine. On a quiet domain a deep page can wait weeks, and with a price change or a closing registration those are precisely the weeks that count. The IndexNow connection reverses the direction: the site itself signals GoogleBot and BingBot as soon as something changes.
For changes that mainly buys lead time. An amended description, a new price, an opened calendar or a reopened section need not wait until somebody happens to pass by. For anything tied to a calendar — events, tourism, hospitality, seasonal ranges — that decides whether you are visible while the demand is there, or only afterwards.
What does change is the question you can ask. No longer: does the search engine know this page at all? But: it knows the page, it has fetched it and still the page is not there — what is holding it back? That second question can be investigated.
Content that has just changed
A new price, an amended opening time, an updated description: here every day you gain counts.
- Offers with an end date
- Pages that carry the season
Addresses that have just become reachable
A lifted block, a removed noindex or a new section that is not yet linked from anywhere.
- After a launch or migration
- After repairing robots.txt
Pages that have not changed
Resubmitting without a change costs daily budget and delivers nothing. The search engine has already formed its judgement.
- Repetition does not speed up a decision
Pages that are being refused
Thin content, a duplicate or a contradictory canonical does not become more acceptable through a notification.
- First the cause, then the submission
From suspicion to status code
It is not the sending but the record-keeping that makes this module usable. For every address it logs when a bot came by, which status it met and which error detail went with it. Three counters run alongside — submitted, found, failed — moving during a job, so you are not waiting in the dark.
The three pieces only work together. The timestamp shows whether anything was fetched, the status betrays whether this is a technical or a content matter, and the error detail points at where the repair sits. That ends a discussion which can linger for years — that Google supposedly bars a page for unfathomable reasons — and starts a task list with owners.
- Submitted, never fetched. The counter records the submission, the log stays empty. Look then at robots.txt, at server load and at possible throttling of the whole domain.
- Fetched, but the server gave way. A 5xx during the visit is rarely a problem of this one page; it presses on throughput for the entire domain.
- Fetched, but nothing was there. Status 404 or 410 points at an outdated row in the sitemap or a link to removed content. Delete it, then, rather than sending again.
- Fetched, but redirected. You submitted a waypoint instead of the destination. From now on submit the final address.
- Fetched, read, refused. Technically everything is in order. The cause then lies in the content, in the canonical reference or in absent internal links.
Forty thousand addresses at a thousand a day
The calculation below exists to make the order of magnitude tangible; it is emphatically not a measurement of an existing project.
An Amsterdam trading platform has 40,000 addresses in its sitemaps, Dutch and English versions together. At 1,000 a day, one complete round takes forty days, and that holds only with the budget fully used and nothing repeated. Add the correction rounds and you are at six to eight weeks. A batch size of 10,000 changes none of that; it only removes the cutting and pasting.
Six weeks is too long for an offering with a shelf life. That argues not against the instrument but in favour of sorting. A plausible breakdown of those same forty thousand looks as follows.
- Roughly 14,500 filter and parameter variants. Sort orders, price ranges, session parameters. This is a canonicalisation question; budget consumed: nil.
- Roughly 9,200 expired items. Past dates, products withdrawn from sale. Those call for a redirect or a clean 410, not for submission.
- Roughly 260 category and landing pages. They carry the lion's share of revenue and fit together into a quarter of one day's budget.
- Roughly 4,400 current offer pages in both languages. Those follow in four to five days and cover the live trade.
- Roughly 11,640 archive and long-tail pages. Those run along afterwards in the background, without a deadline and without pressure.
What looked like forty days shrinks to five for everything that earns money. That gain comes from the sorting, not the technology. A bilingual site adds one criterion: decide explicitly which language goes first. An English product page left out of the index costs an international clientele as much as its Dutch counterpart, yet in a domain-wide average that loss disappears. The visibility reports on the same platform sit in the same account, so that check needs no separate set-up.
Sitemap hygiene and a workable order
A sitemap is not an inventory list but a nomination: with every row you say that this page is worth keeping. If a quarter of the file consists of redirects, error pages and addresses carrying noindex, the standing of the whole drops — including that of the rows which belong there entirely rightly.
Indexable destinations
Only canonical addresses with status 200, written exactly as they finally stand.
- Truthful lastmod, no nightly refresh
- Separate files per language and content type
- An overarching index that stays within three layers
Anything that gets refused anyway
Every address the search engine rejects regardless weakens the point of the file.
- Noindex, blocked addresses, redirects
- Filtered variants and internal search results
- Basket, account area, test hosts, error pages
The working order therefore begins with cleaning, not sending. Sanitise the files first and have them read afterwards, as an upload or a reference, depending on whether they are public already. What comes back is a baseline: how many addresses were found, how many stumbled and which status codes pile up.
You clear that error list before a single day's budget goes into it. Only then does the sorted list go out, load-bearing pages at the top. A week later the log holds enough for three piles: fetched and kept, fetched and refused, never visited. The last pile is technical work, the middle one editorial.
Frequently asked questions
How long does it take before a submitted address is in the index?
There is no hard deadline, and every figure you read somewhere is an estimate. What you do shorten is the route to discovery and to the first fetch. Inclusion itself stays a decision of the search engine. The log at least shows whether that fetch happened, which is already half the answer.
Are a thousand addresses a day enough for a mid-sized webshop?
For the daily rhythm, almost always. Even a busy publishing schedule stays well under a thousand changed addresses per day. It gets tight during a first full inventory, a redesign or a migration — precisely the moments when sorting in advance pays off, because the commercially load-bearing part of a catalogue is usually small.
How do you handle a site serving both Dutch and English?
Both versions are independent addresses and both draw on the same daily budget. Put the files per language under one overarching index — that chain expands without trouble — and then record which language takes precedence. Language detection is not part of the instrument; that choice comes from your own structure and priorities.
What happens if I submit the same address several times?
Then budget goes into nothing. A page that has not changed is not assessed any faster through repetition. A second notification only pays off after a real change: different content, a corrected canonical or a block you have just removed.
Why are there pages in the index that I wanted to exclude?
Nearly always through the confusion between robots.txt and noindex. An address blocked in robots.txt is not fetched, so the noindex inside is never read, and external links can keep the page in the index anyway. Open access, let the noindex be read, and block again afterwards if still needed.
Where visibility actually begins
Indexing is an underestimated bottleneck because it never announces itself as a fault. No alarm sounds and no bar turns red; there are only pages that reach zero impressions. Fail to measure that and you can spend months on copy and links while the blockage sits one phase earlier.
The Indexing Hub does not solve that at the press of a button. It places three things side by side that mean little separately: a throughput limit that forces you to choose, a signal that shortens the route to discovery, and a record that replaces guessing with timestamps and status codes. The boundaries belong to that picture — a thousand addresses a day is finite, so are two parallel jobs, and a notification stays a notification.
That sobriety is what makes the work plannable. Someone who has it on paper that forty thousand addresses need forty days at full budget makes different commitments than someone who hits send and hopes. To find out how many of your pages are genuinely fetched, start with one sitemap round through the Semalt panel and your first indexing job. That first picture is usually more instructive than expected — and on bilingual Amsterdam sites the surprise nearly always sits in the language version nobody was watching.
Hulp Nodig met Uw SEO?
Onze Amsterdam SEO-experts kunnen u helpen deze strategieën te implementeren en uw rankings te verbeteren.
Gratis Consult Aanvragen