Redefine Web
SEO

Website stats checker tools and what they cannot know

A website stats checker estimates other people's traffic rather than measuring it. Where the numbers come from, how far off they run, and when to trust them.

· 15 min read
Website stats checker illustration
Key takeaways
A stats checker models other people's traffic. Only the site owner can measure it, because measurement needs code on the page.
Similarweb states it does not expect its estimates to align exactly with your direct measurement data.
Similarweb publishes a floor of 5,000 visits, below which it displays no data at all.
A blank is not a competitor with no traffic. It is a competitor the tool cannot see.
Calibrate against your own site first, because that is the only place you hold both numbers.

A website stats checker tells you how much traffic somebody else’s site gets, and it does not know. It cannot know. Only the site owner has the measurement, because measurement requires code running on the pages, and these tools have no code on anybody’s pages but their own. What they produce is a model, and the vendors say so in writing when you read far enough into their documentation.

That is not a reason to ignore them. It is a reason to know which questions they answer well and which numbers you should never put in a document with your name on it. This walks through where the figures come from, the published floors below which the tools show nothing at all, why two of them disagree about one domain, and the calibration step that turns an unusable number into a usable one. If you want the broader review first, our guide to running a website audit covers the ground that is not about traffic.

What a website stats checker actually measures

A website stats checker measures nothing about the site you typed in. It measures a sample of internet behavior it has access to, then models the rest, then presents the output as a number with no error bar on it. Every part of that sentence matters and the last part is why people get into trouble.

Website stats checker. An estimator window with a visits chart on one side and a source table on the other, above three metric tiles, the screen a competitor traffic figure is read off.

Compare it with the measured alternative. Analytics on your own site counts events as they happen, from code you installed. Search Console reports what the search engine itself recorded. Both are first party and both are limited to property you control. The moment you want a number about a domain you do not own, measurement stops being available and estimation is the only thing left.

The industry blurs this by using one vocabulary for both. A dashboard showing your sessions beside a competitor’s sessions implies the two numbers were produced the same way. They were not. One is a count and the other is a projection, and putting them in matching tiles is the single most misleading convention in this category.

So the useful habit is to read every competitor figure as carrying a silent word in front of it. Estimated sessions. Estimated keywords. Estimated share. If the sentence still supports your decision once that word is added, the tool has done its job.

It is worth separating this from its neighbors, because the word checker covers several unrelated products that share a name and nothing else. A tool telling you whether a site is safe to visit is doing security work, and our comparison of website checkers covers that end. A tool timing how fast a page loads is doing performance work, covered in our guide to analyzing a website. This article is about one thing only, which is estimating how much traffic a domain receives and how far off that estimate runs.

Why your own numbers and theirs will never match

Point one of these tools at your own site and the figure will differ from your analytics. That is expected behavior rather than a fault, and at least one vendor states it plainly instead of letting you discover it.

Similarweb’s own support documentation says it directly. “Because we are an estimations tool, and we rely on a data collection process that captures billions of digital signals to fuel our powerful algorithms, we don’t expect our estimations to align exactly with your direct measurement data” (support.similarweb.com, Similarweb’s Data Accuracy, read 16 September 2026). The same page sets the expectation it does hold itself to, which is that it provides “a holistic view of the digital world, and expect trend alignment with your direct measurement data”.

Read those two sentences together, because between them they define the product honestly. The vendor is promising that the shape of the line is right, not that the height of it is. A tool that gets direction right and magnitude wrong is genuinely useful for some questions and worthless for others, and knowing which is the whole skill.

The same page adds that “Our data methodology is different than Direct measurement. Hence, there might be a discrepancy when estimating your site.” So when your own site shows a figure you know is wrong, you have not found a bug. You have found the documented behavior of the category, and the sensible reaction is to record how wrong it was and in which direction.

Where the numbers actually come from

Here the picture is more even than this category’s reputation suggests, and the differences between the disclosures are the finding. All three of the major estimators publish a methodology, and they describe two different machines.

Similarweb describes “a unique, multi-dimensional approach” developed over more than ten years, says it provides “statistically representative datasets that preserve variety across countries, industries, user groups, and devices”, and states that it has “been proactive in diversifying our data inputs to be resilient against changes in the market”. That is a real disclosure of shape, even though it stops short of naming the inputs.

Semrush names its inputs and its scale. Its data page says the Traffic and Market toolkit draws on “our panel of over 200 million real, anonymized internet users across more than 190 countries and regions”, built by partnering with “hundreds of clickstream data providers”, and that the clickstream is then processed through an algorithm that “combines various data sources, including our backlink and organic position databases” (semrush.com, Semrush Data and Metrics, read 16 September 2026).

Ahrefs documents a different machine, and it is the one worth understanding. Its help center sets out three steps. It finds “all the keywords for which your target ranks in organic search results”, estimates the traffic from each one “based on its ranking position, monthly search volume, and our estimated CTR for that position”, then adds them up. That is modeled from rankings rather than from observed browsing, which is why the figure covers organic search only. The same page says plainly that these estimates “don’t, and can’t, show you exactly how much organic traffic a website gets”, while adding that “they work incredibly well for comparison” (help.ahrefs.com, What is Organic Traffic in Ahrefs and how do we calculate it, read 16 September 2026).

That contrast matters more than any accuracy claim either vendor makes. A panel-based estimate and a rankings-based estimate are not two attempts at the same measurement, they are two different quantities, and that is most of the reason the numbers disagree. So read which machine produced a figure before you read the figure itself. If you are relying on one of these numbers commercially, it is also fair to ask the vendor how its model treats sites like yours, and the answer tells you something either way.

What you can say across the category with confidence is structural. These tools see a fraction of real browsing behavior and extrapolate from it. That is why coverage is better for large sites than small ones, better in markets where the sample is denser, and better on desktop than on the devices that are harder to observe.

The floors below which the tools show you nothing

This is the most practical published fact in the category and almost nobody knows it. These tools have minimum traffic thresholds, and below them they do not return a bad estimate, they return nothing.

Similarweb publishes its floor. Its support page says “We look at the last available snapshot”, and that if that snapshot “is less than 5000 visits then we won’t display any data”. It publishes a second floor for device-specific features, noting that some of them support desktop data only, so that “if a website has less than 5,000 monthly visits from Desktop” it will have no data to display.

Sit with what that means for competitive research. If you run a local business and your rivals are the same size as you, every one of them may sit under the floor, and the tool will show you blanks or nothing at all. The tools are built for the part of the web with scale, and the small end is not badly covered so much as absent.

The same page explains the N/A you will hit when comparing sites, advising that you “remove the sites which don’t have a sufficient amount of visits, to see data for the other sites”. A blank is not a competitor with no traffic. It is a competitor the tool cannot see, and reading the first as the second is how people conclude a rival is failing when the rival is merely small.

Why two tools disagree about the same domain

Run one domain through three tools and you get three answers, sometimes differing by multiples. The causes are structural rather than accidental, and once you can name them you stop expecting agreement.

Web stats checker. Two estimator windows side by side holding the same monthly visits report, each with its own line and tiles, so one tool's picture of a domain can be set against another's.
  • Different samples. Each vendor observes a different slice of behavior. Two models built on two samples will not converge, however good both are.
  • Different definitions. Visits, sessions, users and pageviews are four different things, and tools headline different ones. Comparing a visit count with a session count is a unit error before it is an accuracy question.
  • Different scopes. Some figures cover organic search only, others cover every channel. A number that looks half the size of another is sometimes measuring half the site’s traffic on purpose.
  • Different index sizes. Where an estimate is built from ranking keywords, the size of the vendor’s keyword index caps what it can attribute. A keyword nobody tracked contributes zero traffic to the model.

That last mechanism has been understood for a long time. When Screaming Frog published a study titled “How Accurate Are Website Traffic Estimators?”, dated 13 June 2016, it set out the prediction that the tools would underestimate, reasoning that “these traffic estimator tools have limited indexes and only track a certain amount of keywords, so can’t possibly expect to completely accurately estimate traffic.” The study compared three tools against analytics data for twenty five sites. Its figures are ten years old and are not quoted here, because the products have been rebuilt several times since, but the mechanism it identified is still the mechanism.

So disagreement between tools is information about the tools, not about the site. If you need a single number, pick one tool and stay in it, exactly as you would with any score. Our comparison of website ranking software covers the same problem on the rank side.

The one number on these tools that may not be an estimate

There is a wrinkle that almost no article about this category mentions, and it changes how you read the output.

Similarweb invites site owners to connect their own analytics to it. Its documentation tells owners who think their site is underestimated to “connect your Google Analytics to Similarweb”, and says this can be done “publicly (all users will see this data) or privately (only you will see this data)”.

Follow the consequence. Some of the figures displayed for some domains are not model output at all, they are measured analytics data the owner chose to publish. Two rows in the same comparison table can therefore have completely different provenance, one modeled and one measured, and nothing in the interface obliges you to notice.

This cuts both ways and both ways are worth knowing. A competitor who has connected their analytics is showing you something close to truth, which is better data than you assumed. A competitor who has not is showing you a model, and if they are the kind of business that would connect it to look impressive, the ones who did not may be systematically different from the ones who did. That is a sampling problem sitting on top of a sampling problem.

Practically, look for a marker on the profile indicating verified or connected data, since platforms offering the option generally badge it somewhere. Where you find one, the figure deserves more weight than the rest of your comparison set and should probably not be corrected by the ratio you derived from your own site, because that correction was built to fix modeling error and this number does not have any. Where you find no marker, assume a model and apply the correction as normal. Recording which rows are which takes a column and it stops you averaging two incompatible kinds of number into one misleading benchmark.

What the estimates are genuinely good for

Having said all that, these tools earn their place. The trick is matching the question to the precision available, and there are three questions they answer well.

The first is relative size. Whether a competitor is roughly your size, several times larger, or an order of magnitude bigger is a question a model can answer confidently, because the error that ruins an absolute figure rarely reverses a ten-fold gap. Most competitive decisions only need this.

The second is direction over time. Because the vendor is promising trend alignment rather than absolute accuracy, a series is more trustworthy than a point. A competitor whose estimated traffic has been climbing for four quarters is doing something, and you do not need the height of the line to know the slope.

The third is discovery. These tools are better at telling you which pages and which topics drive a competitor’s traffic than at telling you how much of it there is. The ranked list is useful even when the volumes beside it are soft, because the ordering survives errors that the magnitudes do not. What that visibility figure is built from is a related question, covered in our explanation of SEO visibility.

How to calibrate a tool before trusting it on anyone else

Here is the step that separates people who use these tools well from people who quote them badly, and it takes about twenty minutes.

You are the only person who can check one of these tools against truth, because you are the only person holding both numbers for your own site. So run your own domain through the tool, pull the same period from your analytics, and write down the ratio. If the tool says sixty percent of your real figure, you have just learned the correction factor to apply to every competitor number it gives you.

Four cautions keep the calibration honest.

  • Match the definitions first. Compare organic sessions with organic sessions, not organic sessions with total visits, or the ratio you derive is a unit conversion rather than an accuracy measure.
  • Use several months, not one. A single month can be unrepresentative on both sides, and you are trying to find a standing bias rather than a fluctuation.
  • Expect the factor to travel badly. The correction you derive holds best for sites like yours. Applying it to a site ten times larger, or one in a different market, is extrapolating from one data point.
  • Redo it periodically. These models are rebuilt. A correction factor from last year is an assumption, not a measurement.

If your own site sits below the published floors, the calibration cannot be done at all, and that is itself the answer. A tool that cannot see you probably cannot see the competitors you care about either. Reading your own measured numbers properly is the prerequisite, and our explanation of SEO analytics covers doing that without fooling yourself.

You can do the whole calibration on free tiers, which is worth knowing before anyone signs anything. Most of these tools expose a limited free view, usually one domain at restricted depth, and that is enough to derive a ratio against your own analytics. What each free tier holds back once you want more than one check is set out in our comparison of free SEO tools. Pay only once you know the tool reads your own site sensibly, because a subscription does not improve a model that cannot see you.

What to ask before quoting a competitor’s traffic

Once a number leaves the tool and enters a document, it stops being an estimate in everybody’s mind and becomes a fact. Four questions before that happens.

  • Which tool, and on what date. Provenance and a date, every time. These models are revised and a figure without a date cannot be reproduced by the person reading it.
  • What is the unit. Visits, sessions, users or pageviews, and organic only or all channels. Half the disagreements in this area are unit mismatches wearing an accuracy costume.
  • Is the subject above the floor. If the site is near the published threshold, the figure is at its least reliable and may be absent for the next domain you check.
  • Would the decision change if the number were half or double. If not, precision was never required and you can stop arguing about it. If so, you need something better than an estimate.

That last question does most of the work. A great deal of effort goes into refining numbers whose exact value would not change anybody’s plan, and the effort would be better spent on the one figure that would. Our breakdown of an SEO report applies the same test to everything else in a document.

What we would do first, and what to ask us

With a competitor set and no budget, we would calibrate before comparing anything. Run our own domain, pull the same window from analytics, write down the ratio, and only then look at anybody else. It costs one session and it converts every later number from a claim into a range.

Then we would use the tools for ordering rather than for magnitude. Which competitors are in the same weight class, which pages earn them the most, which direction each has moved over a year. Those survive the error. The absolute figures do not, and we would not put one in a document without the tool name and the date beside it.

Now point it at us, because agencies quote these numbers in pitches more than anyone. If a figure about your competitors appears in something we send you, ask which tool produced it, on what date, in what unit, and whether we calibrated it against a site we can actually measure. If the answer is that it came from a tool and was typed straight in, treat it exactly as this article says to treat any uncalibrated estimate. That test applies to us as readily as to whoever pitched you last week.

So the first step is free and it is about your own site, not theirs. Run your domain through one of these tools, open your analytics beside it, and find out how far apart they are. If you would rather have somebody look at what the measured numbers are telling you, our free website audit is where to start.

Frequently asked questions

For a site you own, use analytics on your own pages and Search Console, which measure rather than estimate. For a site you do not own, you are limited to estimation tools that model traffic from sampled behavior. The two answers are different in kind, not just in accuracy, and no tool gives you measured numbers for somebody else's domain.

Its published accuracy page describes a multi-dimensional approach developed over more than ten years, producing what it calls statistically representative datasets preserving variety across countries, industries, user groups and devices. It says it captures billions of digital signals to feed its algorithms. The page describes the shape of the method rather than naming every input, so it tells you less about the raw sources than Semrush's data page does.

It is credible for what it claims to be, which is an estimation tool. Its own documentation states it does not expect estimates to align exactly with your direct measurement data and that it expects trend alignment instead. So treat it as reliable for relative size and direction, and unreliable for exact figures, particularly near its published 5,000 visit floor.

You cannot measure it, only estimate it. Use a traffic estimation tool, then calibrate it by running your own site through the same tool and comparing against your analytics to find how far off it runs. Apply that correction to the competitor figure and treat the result as a range. Check the competitor is above the tool's published minimum first.

The useful question is which one you can calibrate rather than which is best in general. Pick one, check it against your own measured numbers, and stay in it, since switching tools changes both the sample and the definitions underneath. If you need a tiebreaker, read the published methodologies, because a panel-based estimate and a rankings-based estimate answer different questions.

For your own site, yes, and properly. Analytics records sessions and conversions from code on your pages, and Search Console reports impressions, clicks and positions from the search engine itself. Both are free and both measure. For other people's sites there is no monitoring, only repeated estimation, so watch the direction of the series rather than any single reading.

For your own site the free first party tools are better than any paid estimator, because they count rather than model. For competitor research most estimation tools offer a limited free view, usually one domain check with restricted depth. Judge a free tier by whether it shows the unit and the date alongside the number, since without both the figure cannot be checked later.

Start with measured data on your own site and establish what normal looks like across a few months. Then add estimated competitor data for relative size and direction only. Match units before comparing anything, note which figures are modeled and which are measured, and write the tool and date beside every estimate. Analysis fails on unit mismatches far more often than on model error.
Found this useful? Share it.
Keep reading
FREE · WRITTEN IN 24 HOURS · NO PITCH

Get your free website audit.

A written report in your inbox within 24 hours, with three fixes you can ship the same week, whether or not you hire us.

WRITTEN IN 24 HOURS · 10,000+ SITES RUN · 300+ CLIENTS SINCE 2021