Experiment Canvas: Design a Test Before You Run It
An experiment canvas is a one-page template that forces you to design a test on paper before you run it: state a falsifiable hypothesis, name the riskiest assumption behind it, pick a method, choose a single metric, and commit to a pass/fail threshold in advance. Design first, then run.
Quick Answer: An experiment canvas captures five decisions in one view — hypothesis, riskiest assumption, test method, metric, and the success threshold you set before you look at the data. It turns a vague hunch into a designed experiment you can pass or fail, so a result means something instead of confirming what you already hoped.
Most founders run experiments the way they check the weather: they launch something, watch a number wiggle, and then decide after the fact whether it was "good." That is not an experiment. It is a vibe with a dashboard attached. The experiment canvas makes you write down what you believe, what would prove you wrong, and where the line sits — all before a single visitor lands or a single interview starts. This guide walks each field, shows a worked example, and explains how the canvas kills the two failure modes that make most "validation" useless: unfalsifiable claims and moving goalposts.
Why design an experiment on paper before running it
Designing the test first is what separates evidence from theater. The single most expensive mistake in early validation is not running the wrong test — it is running a test with no pre-committed definition of success, so any result can be spun as a win. A canvas removes the wiggle room by making you commit before you have any emotional stake in the outcome.
You cannot fool yourself if the goalposts are already planted. Once real data starts arriving, motivated reasoning takes over. A 2% signup rate feels like a triumph on a bad day and a catastrophe on a good one. If you decide the passing bar after seeing the number, you will always find a story where you passed. Writing the threshold down first — while you are still neutral — is the only defense against your own optimism.
A designed test forces a single, answerable question. Vague experiments try to answer "is this idea good?" That question has no data that can settle it. A canvas narrows the scope to one assumption, one method, one metric. You trade a grand unanswerable question for a small answerable one, and small answerable questions are how validation actually compounds.
Writing it down makes the test cheap to critique before it is expensive to run. A hypothesis on paper can be torn apart by a co-founder in five minutes; the same flaw discovered mid-experiment costs you the whole test cycle. The canvas is a design review for your evidence. For where this sits in the broader sequence of validation moves, see the complete guide to startup idea validation, which frames the canvas as one repeatable loop inside a larger decision system.
The discipline traces back to Testing Business Ideas by David Bland and Alex Osterwalder, and the lean-startup lineage before it. What the canvas adds over a loose "let's just try something" is structure: every field is a decision you would otherwise make sloppily or skip.
The experiment canvas fields at a glance
An experiment canvas has five design fields you fill in before the test and one field you complete after. The five up-front fields are the hypothesis, the riskiest assumption, the test method, the metric, and the success threshold; the sixth, the result and decision, gets written once the data is in. The point is sequence: you never touch the result field until the other five are locked.
Here is what each field captures and the failure it prevents when you skip it. Read it as a checklist — a canvas with any row blank is a canvas that will produce an argument, not an answer.
| Field | The question it answers | What goes wrong if you skip it |
|---|---|---|
| Hypothesis | What specific, falsifiable thing do we believe? | You test a feeling no data can confirm or deny |
| Riskiest assumption | Which belief, if wrong, kills the idea fastest? | You run an easy test on a safe assumption and learn nothing that matters |
| Test method | How will we generate evidence, cheaply? | You over-build, or pick a method that can't produce the signal you need |
| Metric | What single number will we watch? | You drown in vanity numbers and cherry-pick the flattering one |
| Success threshold | What result counts as pass vs. fail, decided in advance? | You move the goalposts after seeing the data and always "win" |
| Result & decision | What did we observe, and what will we do about it? | You collect data, feel vaguely informed, and change nothing |
The magic is not in any single box. It is in the fact that the boxes are ordered and interlocking: a weak hypothesis makes the metric meaningless, and a missing threshold makes the whole thing unfalsifiable. Fill them in order, top to bottom, and each field constrains the next.
Field 1: Write a falsifiable hypothesis, not a hope
A hypothesis is a specific belief stated so plainly that data can prove it wrong. If there is no result that would make you abandon the belief, you do not have a hypothesis — you have a wish, and wishes are not testable. The whole canvas rests on getting this field sharp.
The lean-startup community converged on a standard sentence for this, and it is worth using verbatim because its blanks force precision:
We believe that [specific, testable claim].
Ash Maurya and the Strategyzer lineage extend it into a full experiment statement that folds in the method, metric, and threshold, which we will assemble field by field as we go. But the hypothesis itself starts here: one claim about the world you are willing to be wrong about.
A good hypothesis names a subject, an action, and a context. "People want a budgeting app" is not testable — which people, doing what, when? "Freelance designers who bill hourly will connect their bank account to see profitability per client" is testable, because every noun points at something you could go observe.
Strip out adjectives that hide the risk. Words like "easily," "love," and "seamlessly" smuggle in assumptions you have not separated out. "Users will love the onboarding" collapses three claims — they will finish it, understand it, and feel good about it — into one blurry blob. Split them, then pick the one that matters most for field two.
If you cannot imagine the disconfirming result, rewrite it. Before you move on, say out loud: "This hypothesis is wrong if I see ___." If the blank stays empty, the claim is not falsifiable yet and no test will save it.
Field 2: Identify the riskiest assumption behind the idea
The riskiest assumption is the belief that, if false, collapses the whole idea fastest — and the one you have the least evidence for. This field is what stops you from running comfortable tests on things you already basically know while the real killer sits unexamined. You always test the riskiest assumption first.
Every idea rests on a stack of assumptions: that a problem exists, that people will pay, that you can reach them, that they will stick. Most founders instinctively test the assumption that is easiest to check, not the one most likely to sink them.
Rank assumptions on two axes: importance and evidence. An assumption is dangerous when it is both load-bearing (the idea dies if it is false) and unproven (you are running on faith). Plot your assumptions on those two axes and the riskiest one is the top-right corner — critical and unknown. That is your target.
The riskiest assumption is usually about demand, not feasibility. Engineers, in particular, drift toward testing whether they can build something, because that is the familiar risk. But "can we build it" is rarely what kills a startup; "does anyone want it" is. Be honest about which quadrant your idea's real risk lives in.
Testing the riskiest assumption first maximizes learning per dollar. If the scariest assumption survives a cheap test, everything downstream got more likely in one move. If it fails, you just saved yourself from building on sand. Either outcome is high-value, which is exactly why the safe-but-boring assumption is the wrong place to start.
Field 3: Choose a test method that fits the assumption
The test method is the specific action you will take to generate real-world evidence for the riskiest assumption — an interview, a landing page, a concierge run, a fake-door button, a pre-sale. The right method is the cheapest one that can still produce a signal strong enough to move your decision. Fit the method to the assumption, not to your comfort.
The lean-startup sentence continues here:
We believe that [claim]. To verify that, we will [run this test].
Match evidence strength to how much the decision rides on it. A survey is cheap but weak — people say things they will not do. A pre-sale is more work but strong — money is a costly signal. If the assumption is load-bearing and expensive to get wrong, buy stronger evidence; if it is a quick directional check, a lightweight test is fine.
Prefer methods where people spend something. The most reliable signals cost the participant time, money, effort, or reputation. A "yes I'd use this" in an interview is nearly free to give and worth almost nothing. A booked call, a deposit, or a shared email address costs something, and costly actions predict future behavior far better than polite words.
Design the method so it isolates the one assumption. If your landing page tests demand but you also change the price, the audience, and the headline at once, a null result tells you nothing about which variable failed. Hold everything else fixed so the signal you get is attributable to the thing you meant to test.
Field 4: Pick a single metric that reveals the truth
The metric is the one number your test will produce that directly reflects whether the assumption holds. One metric, chosen before the test, tied straight to the hypothesis — not a dashboard of everything you can measure. More metrics do not mean more truth; they mean more room to cherry-pick.
The experiment sentence adds the measurement:
We believe that [claim]. To verify that, we will [run this test], and measure [this metric].
Choose an action metric over an attention metric. Pageviews, impressions, and time-on-site measure whether people noticed. Conversions, sign-ups, pre-orders, and repeat visits measure whether they acted. Only action metrics tell you about demand; attention metrics are the classic vanity numbers that feel good and prove nothing.
Pick the metric before the test, and pick exactly one primary. You may record secondary numbers for context, but name one as the metric that decides the experiment. If you leave the choice open until after you see the data, you will unconsciously promote whichever number looks best. Committing to one primary metric in advance is half of what makes the threshold in field five enforceable.
The metric must be able to fail. If every plausible outcome of your test produces a "good" reading on the metric, the metric is not measuring risk. A real metric has a range where the assumption is clearly disproven. If yours does not, you are back to unfalsifiable.
Field 5: Set the success threshold before you look
The success threshold is the pre-committed line that separates pass from fail — the number you decide before running the test, so the result interprets itself. This is the single most skipped field and the single most important one. Without it, every experiment becomes a Rorschach test where you see whatever you wanted to see.
This is where the full lean-startup statement closes:
We believe that [claim]. To verify that, we will [run this test], and measure [metric]. We are right if [metric hits this threshold].
Write the threshold as a hard number with a comparison. "Good conversion" is not a threshold. "At least 8% of visitors who land on the page enter an email to join the waitlist" is. The number does not have to be scientifically derived — it has to be committed to before you have data, so it is honest rather than convenient.
Set the bar where the decision actually changes. A threshold is only useful if being above or below it leads to different next actions. Ask: "At what result would I keep going, and at what result would I stop or pivot?" The passing line lives between those, high enough that clearing it genuinely de-risks the idea.
Protect the threshold from yourself. The temptation to renegotiate the bar after seeing a near-miss is overwhelming. Tell a co-founder the number in advance, or write it somewhere you cannot quietly edit. A threshold you can move is not a threshold — it is a suggestion, and suggestions do not falsify anything.
Field 6: Record the result and the decision it triggers
The result field captures what actually happened and — critically — the decision that outcome forces. A canvas that ends with "we got 4.2%" is only half done. The point of an experiment is not the number; it is the action the number obligates you to take. This is the field you fill in last, and only after the other five are locked.
This is exactly where Bland and Osterwalder's framework splits the work across two artifacts. The five up-front fields correspond to their Test Card, which plans an experiment as hypothesis → test → metric → success criterion. The result and decision correspond to their Learning Card, which captures the aftermath as hypothesis → observation → learning → decision. Some experiment-canvas layouts merge both onto one page; others keep them as a pair. Either way, the second half is what turns data into a move.
State the observation before the interpretation. Write down the raw number first — "we observed 4.2% waitlist conversion against an 8% threshold" — separately from what you conclude. Keeping the fact distinct from the story about the fact is what keeps the learning honest.
Name the decision explicitly: persevere, pivot, or run a stronger test. A failed threshold does not automatically mean "kill the idea." It might mean the assumption is wrong, the test was too weak, or the audience was off. But you must commit to one next action in writing, or the experiment quietly changes nothing. If you want to see how the planning and learning halves interlock as a pair of cards, the piece on the experiment canvas versus the test card breaks down which artifact owns which field.
A worked example: from hunch to designed test
Watch a vague hunch become a designed experiment by filling the canvas top to bottom. The transformation is the whole value of the tool — the same idea that started as an unfalsifiable feeling ends as a test with a verdict baked in. Here is a founder's raw hunch: "I think busy freelancers would pay for a tool that auto-categorizes their business expenses."
That sentence is untestable as written: no specific audience, no action, no number, no way to be wrong. We run it through the canvas.
Hypothesis. We strip the adjectives and name subject, action, and context: We believe that full-time freelancers who file quarterly taxes will pay a monthly subscription for automatic expense categorization. Now there is a claim with a disconfirming outcome — they might not pay.
Riskiest assumption. The idea rests on several beliefs: the categorization is technically doable, freelancers feel the pain, and they will pay to remove it. Feasibility is low-risk; the team knows they can build it. The dangerous, unproven belief is willingness to pay. That is the target.
Test method. Rather than build the product, we run a pre-sale: a landing page describing the tool with a real "Start 14-day trial — $12/mo" button that collects card details for a founding-member plan. A pre-sale buys strong evidence because it asks for money, not opinions.
Metric. We measure the percentage of qualified visitors — freelancers arriving from a targeted freelance community — who complete the card entry. Not clicks, not signups: card entries, because that is the action that proves willingness to pay.
Success threshold. Before launching, we commit: We are right if at least 5% of qualified visitors enter payment details. Below that, demand is too thin to justify building; at or above it, we proceed to build the trial experience.
Assembled into the standard sentence, the whole design reads as one line: We believe full-time freelancers who file quarterly taxes will pay for automatic expense categorization. To verify, we will run a pre-sale landing page with a real payment step, and measure the share of qualified visitors who enter card details. We are right if at least 5% do.
The result and decision then write themselves: run the page, record the observed percentage against 5%, and commit to build, pivot the offer, or run a stronger test. Notice that nothing here invents a fake conversion number — the verdict is determined by data you have not collected yet, not a figure you hoped for. For a fully populated layout you can copy, see the experiment canvas template and example.
How the canvas prevents vague, unfalsifiable tests
The canvas kills unfalsifiable tests structurally, by refusing to let you leave the disconfirming outcome undefined. Two fields do almost all of this work: the falsifiable hypothesis (field one) and the pre-set threshold (field five). Together they guarantee that some possible result would prove you wrong — which is the entire definition of a real test.
The hypothesis field blocks claims that cannot be checked. By demanding a subject, an action, and a context, it filters out the "people will love it" sentences that have no measurable referent. If you cannot fill the field with something observable, the canvas stops you before you waste a test cycle.
The threshold field blocks post-hoc goalpost-moving. Because the passing bar is committed before data arrives, the result cannot be reinterpreted into a win. This is the difference between science and self-congratulation: a real experiment can embarrass you, and the threshold preserves that possibility.
Compare the two modes side by side. The contrast is stark once the fields are laid against each other, and it is why a canvas produces decisions while a loose test produces debates.
| Dimension | Undesigned "test" | Canvas-designed experiment |
|---|---|---|
| Claim | "Is this a good idea?" | One falsifiable hypothesis with subject, action, context |
| Assumption tested | Whatever is easiest to check | The riskiest, least-proven belief |
| Metric | Whichever number looks best afterward | One primary metric chosen in advance |
| Success defined | After seeing the data | Before running the test |
| Possible to fail? | No — any result can be spun positive | Yes — a defined outcome disproves it |
| Output | A story that confirms the hunch | A decision the data obligates |
The right-hand column is not more rigorous because it is more complicated. It is more rigorous because every field removes a specific way you were about to fool yourself.
Common experiment canvas mistakes to avoid
The most common canvas mistakes all share one root: skipping the discomfort a field is designed to create. The canvas works by forcing hard commitments, so the failure modes are all forms of dodging that commitment. Watch for these five.
Testing a safe assumption instead of the riskiest one. It feels productive to run a test and get a green result, so founders gravitate to assumptions they already half-know are true. A passed test on a low-risk assumption is motion, not progress. Always aim field two at the scariest quadrant.
Leaving the threshold blank until after the data. This is the cardinal sin, because it silently converts every experiment into a confirmation. If field five is empty when the test launches, you are not experimenting — you are collecting decoration for a decision you already made.
Choosing a vanity metric that cannot fail. Impressions and pageviews almost always go "up and to the right," which is exactly why they are useless as experiment metrics. If your metric rises no matter what people actually do, swap it for an action metric that can register a real "no."
Changing more than one variable per test. When a canvas-designed experiment fails, a clean design tells you what failed. If you altered the audience, the price, and the message at once, a null result is uninterpretable. One assumption, one test, one variable.
Filling the result box but skipping the decision. An experiment that ends in a recorded number and no committed next action is a diary entry, not validation. The decision field closes the loop; without it, you learn and then do nothing — indistinguishable from not learning at all.
Key Takeaways
- An experiment canvas is a design tool, not a reporting tool — you fill the five planning fields before running anything, and touch the result field only after the test is done.
- A hypothesis must be falsifiable — name a subject, an action, and a context, and be able to state the exact result that would prove you wrong before you run the test.
- Always test the riskiest assumption first — the belief that is both load-bearing and least-proven, which for most startups is demand, not technical feasibility.
- Choose one primary metric that can fail — an action metric like conversions or pre-orders, not an attention metric like impressions that rises no matter what people do.
- Set the success threshold before you look at any data — a committed pass/fail number is the only defense against reinterpreting a mediocre result as a win.
- The canvas maps onto Strategyzer's Test Card and Learning Card — the five planning fields are the Test Card (hypothesis → test → metric → criterion); the result and decision are the Learning Card (hypothesis → observation → learning → decision).
- Every experiment must end in a committed decision — persevere, pivot, or run a stronger test — because a recorded number with no next action changes nothing.
Frequently Asked Questions
What is an experiment canvas?
An experiment canvas is a one-page template for designing a test before running it. It captures a falsifiable hypothesis, the riskiest assumption, the test method, one primary metric, and a pre-committed success threshold, plus a result-and-decision field completed afterward. Its purpose is to turn a vague hunch into an experiment that can actually pass or fail.
How is an experiment canvas different from a test card?
A test card, from Strategyzer's Testing Business Ideas, plans a single experiment as hypothesis → test → metric → success criterion. An experiment canvas usually covers the same planning ground but often adds explicit space for the riskiest assumption and the post-test result and decision, folding the paired learning card onto the same page rather than keeping it separate.
What makes a good experiment hypothesis?
A good hypothesis is falsifiable: it names a specific subject, a concrete action, and a context, so real data can prove it wrong. "People want this" fails the test; "freelancers who file quarterly taxes will pay $12/mo for expense categorization" passes, because you can imagine the exact result — few of them paying — that would disprove it.
Why set a success threshold before running the test?
Setting the threshold in advance removes your ability to reinterpret a weak result as a win. Once data arrives, motivated reasoning makes any number look defensible. A pre-committed pass/fail line — decided while you are still neutral — is the only reliable protection against moving the goalposts, and it is what makes an experiment genuinely falsifiable.
How many metrics should an experiment canvas track?
One primary metric, chosen before the test and tied directly to the hypothesis. You may log secondary numbers for context, but naming a single deciding metric prevents cherry-picking the most flattering figure after the fact. If every metric you track always trends up regardless of behavior, none of them are measuring the risk you set out to test.
Which assumption should I test first?
Test the riskiest assumption first — the belief that is both critical (the idea collapses if it is false) and unproven (you are running on faith). For most startups that is demand or willingness to pay, not technical feasibility. Testing the scariest assumption first gives the highest learning per dollar, whether it survives the test or fails it.