A page outside the index does not lose to the competition — it is not even in the race. This article shows why discovery, crawling and indexing are three separate phases, what the Indexing Hub in the Semalt panel makes measurable about them, and which limits that instrument deliberately sets.

Almost every SEO plan starts with copy, keywords and links, understandably: that is where the work is visible and discussable. The step before it goes undiscussed. Does the search engine know your page at all, and has it decided to keep it? While that stays open, you are tinkering with something whose existence the search engine does not acknowledge.

Nor is this a matter for large portals alone. An Amsterdam events platform with thousands of performances, a creative-sector webshop with a broad range, a software company with bilingual documentation: in each case addresses appear at a rate nobody tracks by hand. That a backlog exists usually shows only when organic traffic stalls while the editors publish every week.

Scope. This article is not about ranking factors or content strategy. The subject is technical findability: knowing, fetching, keeping — and the instruments that make those three visible.
Foundations · Technical SEO

Three phases often mistaken for one

The most persistent misconception treats indexing as one event that either happens or does not. In reality there are three consecutive phases, each with its own logic and its own kind of blockage. A problem arising in the first phase cannot be repaired in the third, however many submissions you throw at it.

Weeks can pass between those three acts. An address can be known for months without ever being fetched; it can be fetched and still refused. Hold on to that distinction and you ask the only useful question: at which step is it stuck? A broader overview of the modules sits in the description of the rebuilt Semalt panel; here it is purely the indexing part.

What you observePhase where it goes wrongWhat you do
No bot has ever visited the URLDiscoveryCreate an internal link, extend the sitemap, notify actively
Visit logged, but with a server errorCrawlingCheck infrastructure and response times under load
Visit succeeded, page stays out of the indexIndexingBring content, canonical URL and internal signals into line
Only deep pages lag behindDiscovery and crawlingReduce click depth, clean up pagination and filters
New pages only appear after weeksDiscoveryUse active notification instead of waiting for a routine visit
A tip from practice. Establish the phase before you touch a single button in a submission tool. A page with no inbound route at all is a structural matter; a page that is fetched neatly and still refused is a content matter. Those are two different teams and two different agendas.
Diagnosis · Patterns

Six causes that keep repeating

In audits the same causes return with a predictability that is almost comic. The six below cover the great majority of cases, ordered by the phase in which they arise.

CausePhaseHow you recognise itWhat fixes it
Orphan page without inbound linksDiscoveryAnswers with status 200, appears in no sitemap or navigationPlace it in a relevant section and in the sitemap
Excessive click depthDiscoveryTwelve clicks from the homepage, reachable only through paginationFlatter structure, hub pages, better internal linking
Duplicate content through parametersIndexingThe same article via three paths, with sort and filter variantsPick one canonical URL and use it consistently
Contradictory canonicalsIndexingPagination points at page one, targets redirect or are missingPoint the canonical at an existing, indexable URL
robots.txt confused with noindexCrawlingBlocked URL stays in the index, because the noindex is unreadableOpen access and let the noindex be read
Slow or unstable deliveryCrawlingResponse times climb under load, 5xx errors during peaksSort out caching, hosting and error handling

Two of these six deserve a separate warning. The confusion between robots.txt and noindex costs weeks at every launch: a blocked address is not read, so neither is the instruction inside it. And a canonical is a recommendation, not an order; the search engine may pick a different variant than you intended.

Amsterdam adds a complication of its own. A company that publishes its offering in Dutch and in English — the rule here rather than the exception — doubles its stock of addresses. The crawl budget does not double with it. Anyone managing large volumes without control over response times therefore starts not at submission but at the technical foundations.

Instrument · Indexing Hub

The Indexing Hub: the limits are the design

Three functions sit together in the Indexing Hub: submitting addresses, reading sitemaps and recording what the bots then do with them. The most useful part is not the capabilities but the limits, because they determine what is achievable within which period.

1,000
addresses per day per account
10,000
addresses in one batch
3
levels expanded deep
1,000
sitemaps per job
2 / 20
parallel and queued

The smallest number is the most interesting. A thousand addresses a day is generous for a company site of seventy pages, but with a catalogue of tens of thousands it becomes a scarce resource to be allocated, much like a media budget. That is where the gain of the limit sits: it forces a choice that would otherwise be postponed indefinitely.

Indexing Hub · entry point 1

Bulk submission of addresses

For moments when many URLs change at once: a migration, a redesign, a new season.

up to 10,000 URLs per batch
  • One send, no portions. The whole list goes out in one go; splitting it into manageable chunks falls away.
  • The daily budget stays the brake. Handing over 10,000 addresses at once still means ten days of processing. Batch size therefore does not change the lead time.
  • The ordering is the real decision. With a tight budget, the ranking is the only thing you can still steer, and it determines the return of week one.
  • For exceptions, not for the routine. The daily publishing pace of a normal site never comes near this ceiling.
10,000
URLs in one send
1,000
achievable per day
3
counters per running job
Indexing Hub · entry point 2

Sitemap processing

For collections already organised in sitemap indexes, for example by language or by category.

3 levels · 1,000 sitemaps per job
  • Upload or reference. Both an uploaded file and a public address are accepted, so you can also run a set-up through before it goes live.
  • Chains up to three layers deep. If an index points at a further index, that chain is expanded to three layers. Nest deeper than that and you are better off flattening the set-up.
  • A thousand files in one job. That covers the complete split of a large site by language, section or content type in a single round.
  • Two at a time, twenty waiting. Enough for an agency with a handful of Amsterdam clients, provided somebody guards the order: Friday's launch belongs ahead of routine maintenance.
3
nesting levels
1,000
sitemaps per job
2
jobs at a time
20
jobs waiting

Both entry points end in the same observation and share the same daily budget. How the indexing module from Semalt relates to the analytics and campaign parts is set out in the platform overview.

Mechanics · IndexNow

Pushing instead of waiting

Traditionally the search engine fetches when it suits the search engine. On a quiet domain a deep page can wait weeks, and with a price change or a closing registration those are precisely the weeks that count. The IndexNow connection reverses the direction: the site itself signals GoogleBot and BingBot as soon as something changes.

For changes that mainly buys lead time. An amended description, a new price, an opened calendar or a reopened section need not wait until somebody happens to pass by. For anything tied to a calendar — events, tourism, hospitality, seasonal ranges — that decides whether you are visible while the demand is there, or only afterwards.

Submitting a URL is not the same as getting it indexed. A notification is a request for attention, not an agreement: the search engine still decides for itself whether it comes by and whether it keeps the page afterwards. No sending speed exists that gets a thin page, a duplicate or an address carrying noindex accepted after all. Expecting a hundred per cent inclusion rate after a large batch is wishful thinking — and that is not the instrument's fault, but the way indexing simply works.

What does change is the question you can ask. No longer: does the search engine know this page at all? But: it knows the page, it has fetched it and still the page is not there — what is holding it back? That second question can be investigated.

Worth it

Content that has just changed

A new price, an amended opening time, an updated description: here every day you gain counts.

  • Offers with an end date
  • Pages that carry the season
Worth it

Addresses that have just become reachable

A lifted block, a removed noindex or a new section that is not yet linked from anywhere.

  • After a launch or migration
  • After repairing robots.txt
Not worth it

Pages that have not changed

Resubmitting without a change costs daily budget and delivers nothing. The search engine has already formed its judgement.

  • Repetition does not speed up a decision
Not worth it

Pages that are being refused

Thin content, a duplicate or a contradictory canonical does not become more acceptable through a notification.

  • First the cause, then the submission
Evidence · Bot visit log

From suspicion to status code

It is not the sending but the record-keeping that makes this module usable. For every address it logs when a bot came by, which status it met and which error detail went with it. Three counters run alongside — submitted, found, failed — moving during a job, so you are not waiting in the dark.

The three pieces only work together. The timestamp shows whether anything was fetched, the status betrays whether this is a technical or a content matter, and the error detail points at where the repair sits. That ends a discussion which can linger for years — that Google supposedly bars a page for unfathomable reasons — and starts a task list with owners.

Why this makes such a difference. Without a record, non-indexing stays a suspicion that can survive for months. With a timestamp and a status code, a bounded problem is on the table within a week, including the department that solves it.
Planning · Worked example

Forty thousand addresses at a thousand a day

The calculation below exists to make the order of magnitude tangible; it is emphatically not a measurement of an existing project.

An Amsterdam trading platform has 40,000 addresses in its sitemaps, Dutch and English versions together. At 1,000 a day, one complete round takes forty days, and that holds only with the budget fully used and nothing repeated. Add the correction rounds and you are at six to eight weeks. A batch size of 10,000 changes none of that; it only removes the cutting and pasting.

40,000
addresses in the sitemaps
40
days if submitted unsorted
4,660
commercially relevant addresses
5
days for that portion

Six weeks is too long for an offering with a shelf life. That argues not against the instrument but in favour of sorting. A plausible breakdown of those same forty thousand looks as follows.

What looked like forty days shrinks to five for everything that earns money. That gain comes from the sorting, not the technology. A bilingual site adds one criterion: decide explicitly which language goes first. An English product page left out of the index costs an international clientele as much as its Dutch counterpart, yet in a domain-wide average that loss disappears. The visibility reports on the same platform sit in the same account, so that check needs no separate set-up.

Maintenance · Sitemaps

Sitemap hygiene and a workable order

A sitemap is not an inventory list but a nomination: with every row you say that this page is worth keeping. If a quarter of the file consists of redirects, error pages and addresses carrying noindex, the standing of the whole drops — including that of the rows which belong there entirely rightly.

Do include

Indexable destinations

Only canonical addresses with status 200, written exactly as they finally stand.

  • Truthful lastmod, no nightly refresh
  • Separate files per language and content type
  • An overarching index that stays within three layers
Do not include

Anything that gets refused anyway

Every address the search engine rejects regardless weakens the point of the file.

  • Noindex, blocked addresses, redirects
  • Filtered variants and internal search results
  • Basket, account area, test hosts, error pages

The working order therefore begins with cleaning, not sending. Sanitise the files first and have them read afterwards, as an upload or a reference, depending on whether they are public already. What comes back is a baseline: how many addresses were found, how many stumbled and which status codes pile up.

You clear that error list before a single day's budget goes into it. Only then does the sorted list go out, load-bearing pages at the top. A week later the log holds enough for three piles: fetched and kept, fetched and refused, never visited. The last pile is technical work, the middle one editorial.

A monthly check. Setting the file beside the actual state of the site costs half an hour and prevents a list nobody takes seriously any more after a couple of years. The sitemap processing in the panel suits that routine as well as the big first round; related technical analyses are gathered on our blog.

Frequently asked questions

How long does it take before a submitted address is in the index?

There is no hard deadline, and every figure you read somewhere is an estimate. What you do shorten is the route to discovery and to the first fetch. Inclusion itself stays a decision of the search engine. The log at least shows whether that fetch happened, which is already half the answer.

Are a thousand addresses a day enough for a mid-sized webshop?

For the daily rhythm, almost always. Even a busy publishing schedule stays well under a thousand changed addresses per day. It gets tight during a first full inventory, a redesign or a migration — precisely the moments when sorting in advance pays off, because the commercially load-bearing part of a catalogue is usually small.

How do you handle a site serving both Dutch and English?

Both versions are independent addresses and both draw on the same daily budget. Put the files per language under one overarching index — that chain expands without trouble — and then record which language takes precedence. Language detection is not part of the instrument; that choice comes from your own structure and priorities.

What happens if I submit the same address several times?

Then budget goes into nothing. A page that has not changed is not assessed any faster through repetition. A second notification only pays off after a real change: different content, a corrected canonical or a block you have just removed.

Why are there pages in the index that I wanted to exclude?

Nearly always through the confusion between robots.txt and noindex. An address blocked in robots.txt is not fetched, so the noindex inside is never read, and external links can keep the page in the index anyway. Open access, let the noindex be read, and block again afterwards if still needed.

Where visibility actually begins

Indexing is an underestimated bottleneck because it never announces itself as a fault. No alarm sounds and no bar turns red; there are only pages that reach zero impressions. Fail to measure that and you can spend months on copy and links while the blockage sits one phase earlier.

The Indexing Hub does not solve that at the press of a button. It places three things side by side that mean little separately: a throughput limit that forces you to choose, a signal that shortens the route to discovery, and a record that replaces guessing with timestamps and status codes. The boundaries belong to that picture — a thousand addresses a day is finite, so are two parallel jobs, and a notification stays a notification.

That sobriety is what makes the work plannable. Someone who has it on paper that forty thousand addresses need forty days at full budget makes different commitments than someone who hits send and hopes. To find out how many of your pages are genuinely fetched, start with one sitemap round through the Semalt panel and your first indexing job. That first picture is usually more instructive than expected — and on bilingual Amsterdam sites the surprise nearly always sits in the language version nobody was watching.