Where the numbers come from
There are exactly two ways to know how many people visited a website. The site can count them, by running analytics code or reading its own server logs. Or somebody outside can watch a sample of internet users, see how many of them visited, and scale that up. The first is measurement. The second is estimation, and every traffic figure published about a site by a party that does not host it belongs to the second category, no matter how confidently it is formatted.
This is not a criticism of the vendors. Estimation is a legitimate and difficult craft, and a good estimate is genuinely useful for comparison. The problem is presentational: an estimate printed as "1.2M monthly visits" reads exactly like a measurement, and the confidence interval that ought to accompany it is never shown. Once a figure is quoted a second time, the estimate has become a fact.
Panels and their bias
The oldest method is the panel: recruit a population of users, instrument their browsing, and treat them as representative. It is the technique television audience measurement has used for decades, and it has the same structural weakness: the panel is never representative in the way the arithmetic requires.
Panel bias runs in predictable directions. Panellists skew towards people willing to install monitoring software, which skews technical, desktop and young. Corporate networks are under-represented because installation is blocked. Some countries are heavily sampled and others barely at all. For a large consumer site those biases partly cancel; for a niche site with a professional audience they do not, and the estimate can be wrong by a multiple rather than a percentage. Alexa Rank, which was quoted for twenty-six years as though it were a census, was a panel measurement of exactly this kind.
Clickstream and the privacy squeeze
Clickstream data is the modern successor: rather than a recruited panel, aggregate browsing data is purchased from parties that already hold it, such as browser extensions, mobile applications with broad permissions, some internet providers and some security products. The sample is far larger than any panel and its composition is far less knowable, because the vendor buying it usually does not control how it was collected.
Two forces have been shrinking this supply. Regulation has made the collection and resale of browsing data considerably harder to justify, and platform policy has removed a great deal of it at source: extension stores have repeatedly purged products found to be exfiltrating browsing history, and mobile operating systems have tightened the permissions such collection depended on. The visible effect on estimates is a quiet loss of granularity, particularly for smaller sites and for regions with strong privacy regimes.
Modelled search traffic
The third approach does not observe users at all. It observes search engine results. If a vendor knows that a page ranks fourth for a term, and has an estimate of that term's monthly search volume, and a curve describing what share of clicks the fourth position typically receives, it can multiply the three together. Repeat across every term the page ranks for and the sum is an estimate of organic traffic.
Within its own scope this can be quite good, and it has one large advantage: it degrades gracefully. Being wrong about one keyword's volume barely moves a total built from thousands. Outside its scope it is not wrong so much as blind. Direct visitors, newsletter clicks, social referrals, links inside applications, and traffic from search engines the vendor does not model are all invisible. A site whose audience arrives by any route other than the modelled search engine will be reported as having almost no traffic at all.
The failure mode worth remembering
A modelled estimate of zero means "this site does not rank for terms I track", not "this site has no visitors". On a site with a mailing list, an app, or a loyal direct audience, those are completely different statements, and the second one is usually false.
Why estimates disagree
Three methods, three blind spots
| Method | Sees | Misses |
|---|---|---|
| Panel | Whole-session behaviour for a recruited population, across every site they visit. | Anyone unwilling or unable to be instrumented; whole regions; corporate networks. |
| Clickstream | A very large, opaque sample of real browsing, including direct visits. | Composition and provenance. Shrinking with every privacy change. |
| Search model | Organic search demand, reconstructible from public rankings. | Every non-search channel. Reports a strong direct-traffic site as empty. |
Given three methods with disjoint blind spots, disagreement is the expected outcome rather than a sign that one vendor is incompetent. The useful question is never "which number is right" but "which method produced it, and does that method see the channel this site actually uses".
Using an estimate honestly
Three habits make the difference. Quote the source with the figure, always, so a later reader can weigh it. Prefer comparison to absolutes: the same tool run over two sites in the same niche is considerably more trustworthy than either number on its own, because the method's bias applies equally to both. And say "estimated" in the sentence, which costs one word and prevents the figure from hardening into a fact on its second citation.
The wider lesson from this layer is how quickly it disappears. The report on closed measurement tools lists nine services that once supplied numbers of this kind and no longer exist, which means a decade of citations now point at nothing. If you are quoting a figure you would like to still be checkable in five years, name the method as well as the vendor. Vendors close; registry and DNS records do not.

