Startup Hypothesis Templates: Write Testable Assumptions
A startup hypothesis template turns a vague assumption into a specific, testable, falsifiable claim by naming three things: who you mean, what action you expect them to take, and the measurable threshold that counts as a pass. A good hypothesis states, before the test, exactly what result would prove it wrong.
Quick Answer: Write it as one falsifiable sentence: We believe that [specific audience] will [specific action]. We will know we are right if [a measurable outcome clears a threshold] by [date]. If you cannot name the number that would prove you wrong, you are holding an opinion, not a hypothesis.
Every founder is sitting on a pile of assumptions: that the problem is real, that this audience feels it most, that they will pay to make it go away. None of those beliefs are dangerous by themselves. They become dangerous when you build on them without ever checking them — and you cannot check a belief until you have rewritten it as something reality is allowed to contradict. That rewrite is what a hypothesis template does, and this guide gives you three of them plus the rules that make any of them work.
Why vague assumptions cannot be tested
A vague assumption cannot be tested because it never commits to what would prove it wrong. "Freelancers hate invoicing" feels true, survives every possible piece of evidence, and therefore teaches you nothing — no experiment can come back and contradict it.
The difference between an assumption and a hypothesis is falsifiability. An assumption is a belief you are treating as true. A hypothesis is that same belief, rewritten precisely enough that a specific result could disprove it. Eric Ries calls the beliefs a startup bets everything on its leap-of-faith assumptions; the job of a template is to convert one of those leaps into a claim you can actually land on.
Vague claims fail in four predictable ways, and each one maps to a field the template will later force you to fill:
- They hide the number. "A lot of people would use this" can be argued into a win after any result, because "a lot" was never pinned down.
- They blur the audience. "People" is not a group you can recruit, sample, or put an offer in front of.
- They skip the action. "Would be interested" is an attitude, not a behavior you can observe someone actually doing.
- They omit the deadline. With no window, the test never formally ends, so it never forces a decision.
The philosopher Karl Popper gave us the test that still applies: a claim is only meaningfully testable if you can state, in advance, the observation that would refute it. A hypothesis you cannot imagine failing is not a bold bet — it is a sentence dressed up as one. Getting this right is the first move in the wider startup idea validation process, because every experiment downstream inherits the quality of the claim you started with.
The anatomy of a testable hypothesis: subject, action, and threshold
Every testable hypothesis, regardless of which template you use, is built from the same five components. Miss any one and the claim springs a leak — you either cannot run the test or cannot interpret the result once you have it.
The table below breaks down each component, what it pins down, and the specific failure you invite by leaving it out.
| Component | What it pins down | What breaks if you omit it |
|---|---|---|
| Subject (who) | The specific, recruitable group you mean | "People" — you cannot sample, recruit, or target a test |
| Action (what) | The observable behavior that counts as a yes | You measure stated attitude instead of real behavior |
| Threshold (how much) | The pass mark you commit to up front | You rationalize whatever number happens to arrive |
| Timeframe (by when) | The window the test runs in | The test has no end, so it never reaches a verdict |
| Rationale (why) | The underlying belief the result will confirm or kill | A pass or fail teaches you nothing you can reuse |
The takeaway: a strong hypothesis is specific, testable, and falsifiable, and those three qualities are not vibes — they are what you get when all five components are present. Bland and Osterwalder make the same point in Testing Business Ideas, describing a good hypothesis as testable, precise, and discrete: it isolates one distinct thing, so that when the number moves you know exactly which belief moved it. Keep this skeleton in mind as you read the three templates below; each one is just a different, memorable way of guaranteeing every field gets filled.
Template 1 — the "We believe / we will know" hypothesis format
The "We believe / we will know" format pairs a belief with the exact evidence that would confirm it, which is why it is the workhorse of Testing Business Ideas and Lean UX. Its structure is two sentences:
We believe that [X]. We will know we are right if [measurable outcome clears a threshold].
The first sentence carries your belief; the second carries your falsifiability. The "we will know" clause is where the number lives, and it is the half founders are tempted to skip — which is precisely why the template makes it a separate, mandatory sentence. If you cannot complete the second line, you have not finished the hypothesis.
Here is the format filled in for a tool that chases overdue invoices for freelance designers:
We believe that freelance designers who invoice more than five clients a month will sign up for automated payment chasing. We will know we are right if at least a third of the designers we pitch it to start a paid trial within two weeks.
Notice what the sentence now forces on you: a named audience, a concrete action (start a paid trial, not "express interest"), a committed threshold, and a deadline. A useful variant makes the kill line explicit — "we will know we are wrong if" — so you decide in advance what result ends the idea, not just what result saves it. Because this format already separates the claim from its success signal, it drops straight onto an experiment card template, where the belief becomes the hypothesis field and the "we will know" clause becomes the metric and pass mark.
Template 2 — the XYZ hypothesis format for demand
The XYZ hypothesis format compresses a demand belief into a single line: "At least X% of Y will Z." It comes from Alberto Savoia's The Right It, and it earns its place whenever the riskiest thing about your idea is simply whether anyone wants it at all.
Each letter maps to one field you must commit to:
- X is a percentage — the threshold you name before the test, not after.
- Y is your target market — a specific, identifiable, reachable group.
- Z is the concrete action that counts as real engagement: buy, pre-order, sign up, book a call, leave a deposit.
Filled in, a vague belief like "designers would pay to get invoices paid faster" becomes: at least 30% of freelance designers who see our offer will place a paid pre-order. Now it is falsifiable — the number is set, so you cannot quietly move the goalposts once results land.
A market-wide XYZ claim is usually still too big to test this week, so you shrink it. Savoia calls this hypozooming: you narrow the claim to a small, local, immediate slice you can actually put in front of real people now — one community, one channel, one afternoon — while keeping the same percentage and action. If the slice is representative, its result carries a genuine signal about the whole. For the mechanics of choosing a defensible X and writing a number you can stand behind, see the deep dive on the XYZ hypothesis format.
Template 3 — the "If / then / because" experiment hypothesis
The "If / then / because" format is built for change experiments — when you are testing whether a specific intervention causes a specific outcome. Its structure names the action, the expected result, and the reasoning in one line:
If [we make this specific change], then [this measurable thing will happen], because [this belief about the customer].
The "because" clause is the whole reason to use this template. The "if" and "then" alone give you a testable prediction, but the "because" surfaces the underlying assumption — the real thing you are learning about. When the test fails, a bare if/then just tells you the change did not work; the "because" tells you which belief about your customer was wrong, which is the transferable lesson you carry into the next experiment.
A filled example for the same invoicing tool:
If we lead the landing page with "get paid in days, not months" instead of a feature list, then paid trial sign-ups will rise, because designers are motivated more by cash-flow pain than by features.
This format shines for landing-page copy, onboarding tweaks, pricing changes, and growth experiments — anywhere you are altering one variable and watching a metric respond. It is weaker as a first-contact demand test, because it assumes the product and audience already exist; for that earliest "does anyone want this" question, reach for the XYZ format instead.
With all three on the table, the choice comes down to which risk you are attacking right now. This comparison lines them up by best use, what each one forces you to commit to, and where it comes from.
| Template | Best when the risk is | Forces you to commit to | Rooted in |
|---|---|---|---|
| We believe / we will know | Any belief, especially desirability and viability | The signal that would confirm it | Lean UX; Testing Business Ideas |
| XYZ ("At least X% of Y will Z") | Whether anyone actually wants it | A demand percentage and a target market | The Right It |
| If / then / because | Whether a specific change causes an outcome | The intervention and the reason behind it | Growth experimentation |
The takeaway: the templates are not rivals but tools for different jobs — XYZ for raw demand, "if / then / because" for causal change, and "we believe / we will know" as the general-purpose form that wraps around either. Whichever you choose, the discipline is identical: one belief, one predicted outcome, one number you set in advance.
Turning a Lean Canvas assumption into a testable hypothesis
A Lean Canvas is a grid of assumptions, and turning one into a hypothesis means picking the single riskiest box and rewriting it in one of the formats above. Ash Maurya's canvas lays your whole business model out as nine assumptions — problem, customer segment, unique value proposition, solution, channels, revenue, and so on — and every one of them is a belief currently masquerading as a fact.
You cannot test all nine at once, so you prioritize by risk. The move, borrowed from assumptions mapping in Testing Business Ideas, is to sort each assumption on two axes: how important it is (does the business die if it is false?) and how much evidence you currently have (do you actually know, or just hope?). The assumptions that are both critical and unknown are the ones worth a hypothesis right now.
Most canvas assumptions collapse into three classic hypothesis families, and naming the family tells you what to test first:
- Problem hypothesis — does this segment actually have the problem you imagine, badly enough to act? This is where nearly every canvas should start.
- Value hypothesis — once they use your solution, does it deliver enough value that they take a costly action? Ries frames this as testing whether the product truly delivers value in use.
- Growth hypothesis — how will new customers realistically find you? Ries defines this as testing how people discover the product.
Take the canvas box "Problem: freelancers get paid late and it hurts their cash flow." Rewritten as a problem hypothesis in Template 1: We believe that freelance designers rank late payment among their top three business frustrations. We will know we are right if most of the designers we interview raise it unprompted before we mention it. One vague box, now a falsifiable claim. This prioritize-then-rewrite loop is the engine of the broader idea validation workflow, and the canvas is simply the map that tells you which assumption to feed into it next.
Setting success criteria before you run the test
Decide the pass mark before you collect a single data point — otherwise you will unconsciously fit the threshold to whatever number arrives and call it a win. Pre-commitment is the step that makes a hypothesis falsifiable in practice, not just on paper.
Writing the threshold down in advance removes your ability to negotiate with yourself afterward. A result of "18% signed up" is a clear fail against a pre-committed 30% and a triumphant win against a bar you invent at 15% after the fact. Same data, opposite decision — the only thing that changed was when you chose the number.
Four rules keep your criteria honest:
- Name a threshold you can defend, not one designed to pass. If you would build on the result, the number is high enough; if a skeptical co-founder would shrug, it is too low.
- Write the kill criterion too. State the result that makes you stop, not only the one that makes you continue, so a weak signal cannot be spun into "promising."
- Weight the action by its cost to the customer. A paid pre-order clears a lower percentage bar than a free waitlist, because money is far stronger evidence than a click or a survey checkbox.
- Set the timebox. A threshold with no deadline is a test that runs until you get bored, which is not a test.
Record the pass mark, the kill line, and the deadline on your experiment card before you launch, and put them somewhere a future version of you cannot quietly edit. That single act of writing the number down first is the cheapest integrity safeguard in the entire validation playbook.
Common hypothesis-writing mistakes founders make
Most weak hypotheses fail in a handful of recognizable ways, and every one of them is motivated reasoning dressed up as rigor. Knowing the traps in advance is the cheapest way to avoid running a test that was rigged to say yes before it started.
- Bundling multiple variables into one claim. "If we redesign the page and cut the price, sign-ups will rise" cannot tell you which change moved the number. Keep each hypothesis discrete — one variable per test.
- Leaving the threshold blank. A hypothesis with no committed number is an opinion with better grammar. If there is no line the result can fall below, it is not falsifiable.
- Setting the bar after the results arrive. Deciding what counts as success once you have seen the data guarantees a pass and teaches you nothing.
- Choosing a vanity action. A "like," a survey "yes," or a nod costs the person nothing, so it predicts almost nothing. Pick an action with real skin in the game — a payment, a deposit, a booked call.
- Making the audience too broad to recruit. "Small businesses" is not a group you can put an offer in front of this week; "freelance designers who invoice five-plus clients a month" is.
- Writing an unfalsifiable claim. If you cannot picture the specific result that would prove you wrong, rewrite the sentence until you can.
- Picking a threshold engineered to pass. Setting the bar where you already expect to land is not validation; it is a formality you are performing on yourself.
The through-line: a hypothesis exists to give reality a fair chance to say no. Every mistake above is a subtle way of removing that chance while keeping the appearance of a test.
Key Takeaways
- A hypothesis is an assumption rewritten so reality can contradict it. If no result could prove your claim wrong, it is an opinion wearing the costume of a test, and it cannot guide a decision.
- Every testable hypothesis names a subject, an action, a threshold, a timeframe, and a rationale. Whichever template you pick, a missing field is where the claim leaks and the result becomes un-interpretable.
- The "We believe / we will know" format pairs a belief with the evidence that confirms it. The mandatory second sentence is where the number lives, and it drops straight onto an experiment card.
- The XYZ format — "at least X% of Y will Z" — is the fastest way to make a raw demand belief falsifiable. Hypozoom the market-wide claim down to a local slice you can actually test this week.
- The "If / then / because" format is for change experiments, and the "because" is the point. It surfaces the customer belief you are really testing, so a failure teaches you something reusable.
- Set the pass mark, the kill line, and the deadline before you run the test. Pre-committing the number removes your ability to negotiate the result into a win after the fact.
- Weight the action by its cost to the customer. A paid pre-order is worth more than a thousand waitlist emails, because strong evidence comes from what people risk, not what they say.
Frequently Asked Questions
What makes a startup hypothesis testable?
A hypothesis is testable when it is specific, measurable, and falsifiable: it names a recruitable audience, a concrete action you can observe, a threshold you commit to in advance, and a deadline. The decisive test is whether you can state, before running it, the exact result that would prove the claim wrong.
What is the difference between an assumption and a hypothesis?
An assumption is a belief you are treating as true without checking it. A hypothesis is that same belief rewritten precisely enough that a specific experiment could disprove it. The gap between them is falsifiability — a hypothesis names the result that would refute it, and an assumption never does.
What is a good example of a business hypothesis?
"We believe that freelance designers who invoice five-plus clients a month will sign up for automated payment chasing. We will know we are right if at least 30% of those we pitch start a paid trial within two weeks." It names who, what action, what threshold, and by when — all four make it testable.
How many hypotheses should I test at once?
Test one variable per hypothesis, and prioritize ruthlessly by risk. You may run a few separate experiments in parallel if they do not interfere, but each hypothesis should isolate a single belief. Bundling changes into one test means you cannot tell which one moved the metric, which defeats the purpose.
Should I set the success threshold before or after running the test?
Always before. Setting the pass mark in advance is what makes the hypothesis falsifiable in practice; deciding it afterward lets you fit the number to whatever result you got and call it a win. Write the threshold, the kill criterion, and the deadline down before you collect a single data point.