CRYPTO TRAIN
v0.115.16

Writing

What a sitemap cannot do

6 October 2026 · figures measured 6 October 2026

A sitemap is a list of addresses a site would like indexed. It is easy to read as a guarantee, and it is not one. Asked page by page what it had actually done with forty-three of them, Google had never fetched eighteen.

Three answers, not two

The useful thing about asking per address rather than reading a traffic report is that "not indexed" turns out to be two different states:

What Google saysPagesWhat it means
Submitted and indexed20In the index
URL is unknown to Google18Never fetched. Not judged - not seen
Discovered, currently not indexed5Seen, and set aside for now

Those last two are opposite problems. One page is waiting its turn; the other has had its turn and lost. Anything written to improve the second would do nothing for the first.

Why a traffic report cannot tell you which

A search performance report lists the queries pages were found by. It can only contain pages that surfaced in a result, so both of those states are simply absent from it, and absent looks the same either way.

That matters more than it sounds. A site with a few branded impressions and a lot of silence has two completely different diagnoses available - everything indexed and nothing ranking, or half of it never crawled - and the report everybody reaches for first cannot distinguish them.

The page nothing pointed at

Of the eighteen, one was unreachable. It was in the sitemap from the day it shipped, it was listed in the plain-text index this site publishes for language models, and no page anywhere linked to it. Zero inbound links, internal or otherwise.

A crawler finds pages by following links. A sitemap tells it where to look; a link is the reason it bothers. For a domain with no authority yet, the second is doing nearly all the work, and a page with none of it is the last thing in the queue - which in practice means never.

But a link is not a promise either

The tempting conclusion is that the orphan was uncrawled because it was orphaned. The other seventeen uncrawled pages were linked perfectly well - from the footer, from the index that lists them, from the front page in several cases.

So the honest version is narrower and more useful. Being unreachable guarantees a page will not be crawled. Being reachable guarantees nothing at all. One is a bug you can fix in a line; the other is crawl budget, and a four-week-old domain does not have much.

And it does not pick the pages you would

Among the pages built specifically to be found - one per currency pair - the ones that made it into the index were the quiet pairs. The busiest, most-searched pairs were in the uncrawled eighteen.

There is no strategy in that and nothing to copy. It is the order a crawler happened to work in, on a site it has no particular reason to hurry over. Worth knowing mostly so that a month of flat numbers is read as what it is rather than as a verdict on the writing.

What to actually do about it

Check that every page you publish is reachable by following links from another one. It is a cheap thing to get wrong - a page gets written, added to the sitemap, and the link that would have introduced it is the step nobody remembers - and it is cheaper still to test for.

Beyond that, patience, and the ordinary things: pages worth linking to from elsewhere, and enough of a reason for a crawler to come back. The splitting and providers pages exist on that theory: they are useful to somebody who will never run a transfer here, which is the kind of page that earns a link.

Method

Every address in this site's sitemap was inspected individually against Google's own record for this property on the date above, one request per address, and the counts are theirs rather than inferred from traffic. Reachability was measured separately, by looking for a link to each published page from anywhere else in the source.

Plan an exchange

Every figure above can be checked against the providers yourself — no account, no sign-up, and nothing is held for you at any point.