In this Article
Market sizing figures circulate as facts. A number appears in a report, gets quoted in a deck, repeated in an article, and cited back into the next report, by which point nobody can say where it came from or what it counted.
This guide explains the three methods that produce these numbers, why two reputable reports on the same market can differ by a factor of three, which sources are genuinely reliable, and how to build an estimate you can actually defend.
Key Facts
- Most market size figures are estimates built on estimates. A single cited number typically rests on a chain of assumptions that the summary never shows.
- Three estimation methods produce almost all of them, and knowing which was used tells you how to interrogate the result.
- Reports on the same market often differ by multiples, usually because they define the market differently rather than because one is wrong.
- Forecast CAGRs are the least reliable component and the most quoted, because they are the part that sells the report.
- A bottom-up estimate you built is worth more than a top-down number you bought, because you can defend the assumptions.
How are market size figures produced?
Reliable market research data sources feed three estimation methods, often combined. We call it the 3-part sizing model.
| Method | How it works | Where it breaks |
|---|---|---|
| 1. Top-down | Start from a large published total and apply share assumptions | Every assumption multiplies; small errors compound fast |
| 2. Bottom-up | Count units or customers and multiply by price | Needs real counts, which is work; usually the most defensible |
| 3. Supply-side | Sum known vendor revenues and extrapolate the rest | Private company revenue is estimated, often from headcount |
The tell in any report is whether the method is stated. A figure presented without its method cannot be evaluated, and in most published summaries the method is described in a sentence that omits the assumptions that actually drive the result.
Why do reports disagree so much?
Mostly definition, not error.
Market boundaries differ. Does a proxy market include VPN consumer subscriptions? Does an analytics market include the services revenue around it? Each choice can double or halve the total, and both definitions can be legitimate.
Geography and segment differ. A global figure and a North American figure get quoted interchangeably once they escape the report.
Timing differs. A 2026 figure may be a forecast made in 2024 rather than a measurement, and the distinction disappears in the citation.
Private revenue is estimated. Supply-side sizing depends on guessing revenue for companies that publish nothing, commonly from headcount multiples that vary enormously by business model.
The practical response when you see two very different numbers is not to pick the credible one but to find the definitions. Usually they are measuring different things, and knowing which one matches your question is the whole exercise.
Which sources are actually reliable?
| Source | Use it when | Avoid when |
|---|---|---|
| Government statistics | You need defensible baselines: industry counts, trade, employment | You need a market defined the way vendors define it |
| Public company filings | Segment revenue for listed players is disclosed | The market is mostly private companies |
| Industry associations | They collect from members directly | Membership is partial, which biases the total downward |
| Your own data | Conversion rates and price points anchor a bottom-up estimate | Your segment is unrepresentative of the whole market |
| Commercial reports | For structure, vendor lists and vocabulary | You need the headline number to be load-bearing |
The last row is the one worth internalising. Commercial reports are genuinely useful for understanding how a market is segmented and who the players are. Their headline totals are the least reliable part and the part everyone quotes.
How do you build an estimate you can defend?
Bottom-up, with every assumption visible.
- Count something real. Companies in a segment, listings, job postings, app installs, public customer logos: an observable count beats an inherited percentage.
- Anchor price from evidence, using published pricing or your own average contract value rather than a market-wide assumption.
- Write the assumptions as a list with a range on each, and compute a low, mid and high case. A single number implies precision you do not have.
- Cross-check against a top-down figure. Agreement within an order of magnitude is reassurance; a large gap tells you a definition is different, which is information rather than a problem.
The observable counts in the first bullet usually come from public sources that render differently by country: business directories, job boards, marketplaces and company sites. DataImpulse residential pins that collection to a country at $1 per GB across 195 countries, which matters because a count gathered from one market and presented as global is exactly the error this method exists to avoid.
Related: company data APIs, data enrichment sources.
Frequently Asked Questions
Where do market size numbers come from?
From three estimation methods: top-down, applying share assumptions to a large published total; bottom-up, counting units or customers and multiplying by price; and supply-side, summing vendor revenues and extrapolating. Most published figures combine them and state the method only in passing.
Why do market reports disagree by so much?
Usually because they define the market differently, not because one is wrong. Boundary choices about what to include, geographic scope, forecast versus measured timing, and estimated revenue for private companies can each change a total by a factor of two or more.
Are commercial market reports worth buying?
For segmentation, vendor landscape and vocabulary, often yes. For a headline number that will carry weight in a decision, no: that figure is the least reliable part of the report and the most frequently quoted out of its definition.
How do I build my own market estimate?
Bottom-up. Count something observable, anchor price from published evidence or your own contract values, list every assumption with a range, and produce low, mid and high cases. Then cross-check against a top-down figure to see whether the definitions match.
Which sources are most reliable for market data?
Government statistics and public company filings, because they are defensible and auditable, followed by industry associations with the caveat that partial membership biases totals downward. Your own conversion and pricing data is the strongest anchor for a bottom-up estimate.
Count the same way in every market
Directories, job boards and marketplaces show different results per country, and a count from one market presented as global is the error bottom-up sizing exists to avoid. DataImpulse residential proxies pin collection to a country at $1 per GB across 195 countries. Create an account and count consistently.
Related: company data APIs · data enrichment · proxies for web scraping.
Last updated: September 17, 2026.

State/City/Zip/ASN Targeting 



