How to Size a Market When There's No Data
When no analyst report covers your market, you build the number yourself. Decompose the market into a chain of small factors you can estimate — Fermi-style — anchor each one with a proxy or an analogous market, then triangulate two or three independent builds. A defensible estimate with stated assumptions beats a borrowed big number every time.
Quick Answer: No data doesn't mean no estimate. Split the market into a chain of factors you can reason about, anchor each with a proxy or an analogous market, and triangulate a few independent builds. Show every assumption — a number a skeptic can argue with beats a headline they can only distrust.
Founders hit this wall constantly: you're sizing a brand-new category, you search for a market report, and nothing credible comes back. That is not a research failure. The report doesn't exist because the category doesn't yet — and pasting a vaguely related "$XX billion" figure from a press release fools no one who reads decks for a living.
So change the goal. Market sizing with no data isn't a lookup; it's a construction. You're not finding a number someone already calculated — you're building one from parts you can defend, and making your reasoning visible enough that a skeptic argues with a single assumption instead of dismissing the whole slide.
This is the hard-mode version of the broader startup market sizing process, for the exact case where the usual reports don't exist. Four moves get you there: frame the customer, borrow from analogs and proxies, stack your assumptions into a build, and make each one falsifiable.
Prerequisites: Frame the Customer and the Job Before You Estimate
Before you estimate a single number, pin down who the customer is and what job they hire your product to do — because every factor you're about to estimate multiplies off that definition. A fuzzy customer corrupts the whole chain.
Name the buyer tightly enough to count. "Students" is not a market you can size; "final-year undergraduates who already pay for exam-prep tools" is. The narrower the definition, the easier it is to find a real anchor for how many exist — and the harder it is for anyone to accuse you of hand-waving.
Define the job, not the product. What outcome is the customer buying — the thing they would hire any solution for? Framing it as a job ("help me revise efficiently under time pressure") rather than a feature list is what lets you spot analogous markets later, because different products often serve the same underlying job.
Fix the unit and the time window. Decide what one customer is (the account you would invoice) and measure spend per year. Every factor you estimate feeds the same skeleton as a bottom-up TAM calculation — potential customers times annual spend — you just estimate the inputs instead of looking them up.
Write down what you're estimating as one sentence. "Annual revenue if every [tightly defined buyer] paid [annual price] for [the job]." That sentence is the target the rest of the work fills in, and it keeps you from silently sizing a different market halfway through.
Use Analogous Markets and Proxies as Anchor Numbers
When your market has no numbers, borrow them: an analogous market or a measurable proxy gives you an anchor to build from instead of a blank page. Both swap a wild guess for something traceable to a real source.
Pick an Analog That Shares the Buyer, the Job, or the Adoption Curve
An analogous market is an existing, measured market that resembles yours closely enough to lend its figures as a starting point. The best analog shares at least one of three things with you: the same buyer, the same job-to-be-done, or the same adoption pattern.
Borrow the rate, not the headline total. What you want from an analog is a ratio you can reapply — what share of the base actually adopts, what a customer spends per year, how deep penetration goes at maturity — not its absolute market size, which describes its world, not yours.
Reject flattering analogs. If the only thing your category shares with a giant market is a buzzword, it is a vanity analog and it will inflate your number. A humble, structurally similar market is worth more than a huge, superficially similar one.
Use Proxy Metrics When Even an Analog Is Thin
A proxy is something you can measure that stands in for something you can't. You cannot directly count "people frustrated with clunky revision tools," but you can count downloads of competing apps, search volume for the problem, or members of a relevant association.
Pull those stand-ins from free public data sources — census and labor statistics, trade-body and association reports, and the filings of public companies that already sell into the space. The table maps common proxies to what each one approximates and where to find it.
| Proxy you can measure | What it stands in for | Free source to pull it from |
|---|---|---|
| Members of a trade or professional body | Size of a defined occupation or buyer group | Association directories and annual reports |
| Downloads or review counts of rival apps | Users tolerating today's tools | App-store listings, public review pages |
| Search volume for the problem | Latent demand and awareness | Free keyword and trends tools |
| Job postings naming the task | Firms with the pain, hiring against it | Public job boards, labor-statistics releases |
| Segment revenue of a public comparable | Money already flowing to the job | Company annual reports and regulatory filings |
| Census counts for the occupation or firm type | The universe you will narrow from | National statistical and census agencies |
Takeaway: No proxy is exact, and that is fine — you are not seeking precision, you are seeking an anchor with a traceable origin. Three mediocre proxies that roughly agree are far more convincing than one confident guess with no source behind it.
Build a Stacked-Assumption Estimate With the Fermi Method
A stacked-assumption estimate breaks the market into a chain of factors you can each estimate, then multiplies them — so a number you can't look up becomes several smaller ones you can. This is Fermi estimation, named after the physicist who taught students to approximate impossible-seeming quantities (famously, the piano tuners in a city) by chaining rough guesses whose errors tend to cancel out.
The method is a funnel. Start from a broad, sourceable universe and multiply by a series of narrowing rates until you reach paying customers, then multiply by annual price. Every step gets an explicit assumption written next to it.
Here is an illustrative walk-through for a hypothetical scheduling tool sold to independent music teachers. Every figure below is invented to demonstrate the method — these are not researched numbers, and you should never quote a teaching figure like this as a real market size.
| Factor (all figures illustrative) | Assumption being made | Where the anchor comes from |
|---|---|---|
| 200,000 private music teachers nationally | The countable universe | Teachers' association + census occupation code |
| × 40% teach enough to treat it as income | Only real earners buy tools | Membership survey, analog market |
| × 50% would pay for scheduling software | Half are willing to pay | Rival app user counts (proxy) |
| = 40,000 potential customers | The addressable base | — |
| × $120 assumed annual price | What one would accept per year | Competitor pricing pages |
Multiply the chain and the illustrative estimate is 40,000 × $120 = $4.8M — a made-up teaching figure, not a claim about any real market. What matters is not the $4.8M; it is that the number is now a stack of four visible assumptions, each of which a reader can challenge one at a time.
Triangulate: Estimate the Same Market a Second Way
Triangulation means reaching the same number by an independent route and checking whether the two agree. One build can be wrong in a way you cannot see; two builds that disagree expose the bad assumption for you.
Cross-check with an analog. Suppose the closest comparable — an established booking app for a similar solo profession — reports roughly 8,000 paying users in the same country (again, illustrative). If that product has reached perhaps a fifth of its own base, it implies an order of magnitude near 40,000 addressable users — the same ballpark as the bottom-up chain.
Read the agreement, not the digits. When two independent estimates land in the same order of magnitude, confidence rises. When they diverge tenfold, you have not failed — you have found the single assumption most worth investigating next.
Make Every Assumption Falsifiable So the Estimate Improves
An estimate is only defensible if each assumption is written as a claim someone could prove wrong and cheaply test. That is the line between a guess and a hypothesis — and it is what makes a no-data estimate credible rather than merely hopeful.
List the assumptions as a ledger. For each factor, write the number, your confidence in it, and the cheapest way to firm it up. "50% would pay — low confidence — test with 20 customer interviews" is honest; a lone bold number pretending to precision is not.
Attack the most sensitive assumption first. Ask which input, if it were wrong by half, would move the answer most. In the music-teacher chain, the willingness-to-pay rate swings the total more than the census count does, so that is where your first week of evidence-gathering should go.
Report a range, not false precision. The honest output of market sizing with no data is an order of magnitude — "low single-digit millions," not "$4,812,000." Stating a range signals you understand the uncertainty instead of hiding it behind decimal places.
Keep the estimate alive. Store the assumption list somewhere you will revisit — a spreadsheet, a doc, or a founder workspace like Edmired — because every real deal, interview, and sign-up should sharpen a specific assumption. A borrowed big number can only go stale; a stacked estimate gets sharper, not louder, as evidence arrives.
Key Takeaways
- No data doesn't mean no estimate. For a brand-new category the report simply does not exist yet, so the job shifts from looking a number up to constructing one you can defend line by line.
- Frame the customer and the job first. A tightly defined buyer and a clear job-to-be-done make the market countable and help you spot the right analogs — every downstream factor multiplies off that definition.
- Anchor with analogs and proxies, not guesses. Borrow adoption rates and spend from structurally similar markets, and use measurable proxies — association counts, app downloads, search volume — pulled from free public data.
- Build bottom-up with the Fermi method. Chain a sourceable universe through a few narrowing rates to paying customers times annual price, writing an explicit assumption beside every single step.
- Triangulate to test the number. Reach the same market a second way; agreement in order of magnitude builds confidence, and divergence points straight at your weakest assumption.
- Make every assumption falsifiable. Log each input with a confidence level and the cheapest test, attack the most sensitive one first, and report a range instead of false precision.
- A defensible estimate beats a borrowed big number. Reviewers trust visible reasoning they can argue with, not an impressive figure with no provenance — the assumptions are the deliverable, the digits stay provisional.
Frequently Asked Questions
Can You Estimate Market Size With No Data?
Yes. When no report exists, you construct the estimate instead of looking it up: break the market into a chain of factors, anchor each with a proxy or an analogous market, and multiply. The result is an order-of-magnitude figure backed by stated assumptions — which is far more defensible than a borrowed headline number nobody can trace.
What Is a Proxy in Market Sizing?
A proxy is a measurable quantity that stands in for one you cannot measure directly. If you cannot count frustrated buyers, you count downloads of rival apps, members of a relevant association, or search volume for the problem instead. Good proxies come from free public data, so each one stays traceable to a source rather than invented.
How Accurate Does a No-Data Market Estimate Need to Be?
Only accurate to the right order of magnitude. The goal is not a precise figure — it is a defensible one that shows whether the opportunity is worth pursuing and rests on assumptions you can test. A transparent range like "low single-digit millions" signals rigor; false precision like "$4.8M" invites doubt you cannot yet answer.