Testing Business Ideas: The Experiment Playbook
Testing Business Ideas by David Bland and Alexander Osterwalder is a method for de-risking ideas with cheap experiments instead of business plans. You extract the assumptions your idea rests on, map them by importance and evidence, then run the smallest experiment that produces evidence strong enough to justify your next move: persevere, pivot, or kill.
Quick Answer: Don't defend the idea — test the assumptions underneath it. Map them by importance and evidence, attack the important-but-unproven ones first, and design each experiment to buy the strongest evidence you can afford before you commit real money.
A business plan is a stack of assumptions written in confident prose. Testing Business Ideas — the 2019 field guide from David Bland and Alexander Osterwalder, part of the Strategyzer series behind the Business Model Canvas — treats that stack as the actual work. Instead of arguing about whether the idea is good, you isolate each assumption it depends on and buy evidence about the risky ones, one cheap experiment at a time.
This playbook walks the full method: why experiments beat plans, how to map and prioritize assumptions, the three risks every idea carries, the test card and learning card that keep experiments honest, how to choose from the book's library of 44 experiments, and how to sequence weak signals into strong proof.
Why Cheap Experiments Beat Business Plans
Cheap experiments beat business plans because a plan is untested opinion formatted to look like a decision. It reads as rigorous, but every number in it is a guess wearing a spreadsheet. An experiment does the opposite: it produces evidence you didn't have before, at a fraction of the cost of building the whole product and waiting to see who shows up.
Bland and Osterwalder frame idea development as the R&D of new business models — a discipline of reducing uncertainty over time. The most expensive way to discover you were wrong is to launch the finished product. The cheapest is to run a small test this week that could prove one assumption false before it costs you a year.
This reframes the founder's job. You are not trying to be right; you are trying to find out where you're wrong while it's still cheap to be wrong. Three shifts follow from that:
- From selling to learning. The goal of an early test is not traction. It is a truthful read on whether an assumption holds.
- From opinion to evidence. "Customers will love this" is a hypothesis, not a fact, until behavior says otherwise.
- From big bets to small ones. You break one large, unfalsifiable bet ("this business will work") into many small, testable ones.
The asymmetry is what makes the trade obvious once you see it. A customer interview costs an afternoon; discovering the same insight after launch costs a product, a runway, and often a team. Bland and Osterwalder describe this as buying information — the question is never whether a test is perfect, but whether it's the cheapest thing that meaningfully shifts your confidence in a risky assumption. Framed that way, skipping the test to "just build it" is rarely the fast path. It's the most expensive way to stay uncertain.
That's also why the method fits corporate innovation teams and solo founders equally well. A new venture inside a large company faces the same three risks as a garage startup; it just has more to lose by defending a plan through a stage gate instead of testing it. In both settings, the unit of progress is the same: one assumption meaningfully de-risked.
If you haven't yet inventoried what your idea actually assumes, start upstream with the complete guide to startup idea validation — a beautifully run experiment aimed at the wrong assumption teaches you nothing that matters. The method below assumes you know which guesses are load-bearing.
The Testing Business Ideas Method at a Glance
The method is a repeatable loop, not a one-time checklist: shape the idea into something testable, extract its assumptions, prioritize them, test the riskiest, capture what you learned, and decide. Bland and Osterwalder call this the Design-Test loop, and the whole point is to travel around it quickly and cheaply, many times.
Here is how each step maps to a concrete Strategyzer tool and the output it should produce:
| Step | Tool | Output |
|---|---|---|
| Shape the idea | Business Model Canvas & Value Proposition Canvas | An explicit, testable business model |
| Extract assumptions | Assumptions list (desirability, feasibility, viability) | The hypotheses the idea depends on |
| Prioritize | Assumptions Map (importance vs. evidence) | Your riskiest, least-proven assumptions |
| Design the test | Test Card | An experiment with a pre-set pass/fail bar |
| Run and capture | Learning Card | Evidence plus a documented insight |
| Decide | Design-Test loop | Persevere, pivot, or kill — then repeat |
The takeaway: every row hands its output to the next row, so a weak first step quietly corrupts everything downstream. A vague business model produces vague assumptions, which produce experiments that can't fail cleanly. Get the shaping right and the rest of the loop has something real to bite on.
Two properties make the loop work. First, it's iterative — you don't march through the six steps once and graduate; you lap them, and each lap should cost less and teach more than another round of plan revisions would. Second, it's evidence-driven at every handoff — an assumption only advances to the next step when a result earns it that promotion. Teams that treat the loop as a linear checklist tend to run one experiment, declare victory, and jump straight to building, which quietly reintroduces exactly the risk the method exists to remove.
Assumptions Mapping: Importance vs. Evidence
Assumptions mapping is how you decide what to test first, and the answer is rarely the assumption you're most excited about. You plot each assumption on two axes — how important it is to the idea working, and how much evidence you already have for it — and the quadrant that demands attention is high importance, low evidence.
Those are your leap-of-faith assumptions: essential to the idea and still unproven. They belong at the front of the test queue because that is where uncertainty is both highest and most dangerous. An assumption that's important but well-evidenced doesn't need a test yet; one that's unproven but unimportant doesn't deserve your week.
The map sorts your list into four practical piles:
- Important + no evidence: test these first. If one collapses, the idea changes shape or dies.
- Important + have evidence: monitor, but don't spend scarce experiment budget re-proving what you already know.
- Unimportant + no evidence: park them. Uncertainty here is cheap to be wrong about.
- Unimportant + have evidence: ignore. This is where founders love to hide, because the tests are easy and comfortable.
A quick example shows why the map matters. Suppose the idea is a scheduling tool for small physiotherapy clinics. The assumptions might include: clinics lose real money to no-shows (desirability), owners will switch away from their current calendar (desirability), the tool can integrate with the practice-management systems clinics already run (feasibility), and clinics will pay a monthly fee rather than keep using a free calendar (viability). Written out, some are obviously riskier than others. "Owners dislike no-shows" is important but well-evidenced — most will tell you unprompted. "Owners will switch and pay" is just as important and almost entirely unproven. The map sends you at the second one first.
The discipline is resisting the pull toward comfortable tests. It feels productive to validate the thing you already believe; it is far more useful to attack the thing that could kill you. For the full mechanics of scoring and ordering, assumptions mapping and finding your riskiest hypothesis goes deeper than the quadrant sketch — but the principle is simple enough to run on a whiteboard today. Whether you keep the map on that whiteboard, in a spreadsheet, or in a purpose-built tool like Edmired, the rule holds: one assumption, one test, one threshold set in advance.
Desirability, Feasibility, and Viability: The Three Business Risks
Every assumption you extract falls into one of three risk categories, and a durable business has to clear all three. Bland and Osterwalder organize the entire book around this trio: desirability (do customers want it?), feasibility (can we build and deliver it?), and viability (can we make money doing it?). Ignore any one and the other two won't save you.
The value of naming the risk type is that it points you at the right kind of experiment. A desirability doubt and a viability doubt call for completely different tests, and confusing them is how teams "validate" the wrong thing.
The table below maps each risk to the question it answers, what failure looks like, and the experiments that probe it — all qualitative, because the point is direction, not a false-precision score:
| Risk | The question it answers | What failure looks like | Example experiments |
|---|---|---|---|
| Desirability | Do customers want this? | You build it and nobody shows up | Customer interviews, landing page, fake-door test |
| Feasibility | Can we build and deliver it? | You can't deliver at quality, cost, or scale | Single-feature MVP, technical spike, partner interviews |
| Viability | Can we make money doing it? | Customers won't pay enough, or costs outrun revenue | Pre-sale, letter of intent, mock sale, pricing test |
The takeaway: most first-time founders over-test feasibility (the part engineers enjoy) and under-test desirability and viability (the parts that actually kill companies). Balance your experiment budget across all three, and start wherever the assumptions map says the important-but-unproven risk lives — not wherever the building is most fun.
A useful habit is to phrase each risk as a falsifiable sentence. Desirability becomes "clinics will adopt this over their current calendar." Viability becomes "enough of them will pay our target price to cover the cost of serving them." Feasibility becomes "we can integrate with the systems they already run." Phrased that way, each risk names the experiment that could break it — which is the whole reason to sort assumptions by type before you design a single test. The category isn't bureaucracy; it's a pointer to the right tool.
The Test Card and Learning Card: Designing and Capturing Experiments
The test card and learning card are the two paper tools that keep an experiment from drifting into wishful thinking. The test card is filled out before you run — it forces you to commit to a pass/fail bar in advance. The learning card is filled out after — it forces you to convert raw results into a decision. Together they close the loop between "we ran something" and "we learned something and acted."
The test card has four fields: the hypothesis you're testing, the test you'll run, the metric you'll measure, and the criteria that would prove you right. The criteria field is the one that matters most, because writing your success threshold before you see the data is the single best defense against rationalizing whatever you get.
The learning card mirrors it: the hypothesis you believed, the observation you actually made, the learning you drew from it, and the decision you'll now take. This side is where evidence becomes action instead of a slide nobody revisits.
Here is how the two cards line up, field by field:
| Element | Test Card (before the experiment) | Learning Card (after the experiment) |
|---|---|---|
| Hypothesis | "We believe that [assumption]." | "We believed that [assumption]." |
| Action | Test: "To verify, we will [experiment]." | Observation: "We observed [result]." |
| Measure | Metric: "And measure [signal]." | Learning: "From that we learned [insight]." |
| Bar / next move | Criteria: "We're right if [threshold]." | Decision: "Therefore, we will [action]." |
A filled-in pair makes the discipline concrete. A test card might read: We believe a meaningful share of clinic owners who reach our pricing page will start a trial. To verify, we'll drive a targeted audience to a landing page and measure trial sign-ups. We're right if sign-ups clear the threshold we set before launch. Suppose the result lands well short. The learning card then reads: We believed enough owners would start a trial. We observed far fewer. We learned the value proposition isn't landing with this segment as written. Therefore, we'll rewrite the offer around no-show recovery and retest before building anything. Notice that a disappointing result still produced a clear next move — that is the loop working, not failing.
The takeaway: the criteria-then-decision spine is what turns activity into learning. Skip the criteria field and every result becomes "encouraging"; skip the decision field and every learning becomes a note nobody acts on. The cards are deliberately boring on purpose — their job is to remove your ego from the read.
Choosing From the Experiment Library: 44 Experiments by Cost and Evidence
The book's experiment library holds 44 experiments, and choosing among them comes down to two questions: which risk am I testing, and how strong does the evidence need to be? Bland and Osterwalder rate every experiment on cost, evidence strength, setup time, and run time, then sort them into two phases — Discovery and Validation — so you can match the tool to the moment.
Those four ratings are what make the library usable rather than just a list. Cost and setup time tell you what you're risking to run the test; run time tells you how long until you have an answer; and evidence strength tells you how much to trust that answer once it arrives. A cheap test with a fast turnaround and weak evidence is perfect early, when you're wrong about a lot and need to fail fast. An expensive, slow, strong-evidence test only pays off once you've narrowed the field enough that a real bet is the next honest move.
Discovery experiments are cheap, fast, and produce weaker evidence. They're for finding direction early, when you know little: customer interviews, online ads, fake-door or feature-stub tests, explainer videos, storyboards, and paper prototypes. You run them to point yourself in the right general direction, not to bet the company.
Validation experiments cost more, run longer, and produce stronger evidence. They're for confirming a direction once discovery has narrowed it: single-feature MVPs, Wizard of Oz tests, concierge MVPs, letters of intent, pre-sales, and split tests. You earn the right to run these; you don't start here.
The book leans on recognizable cases to make the categories concrete. Dropbox's early explainer video is the canonical discovery move: rather than build the hard file-synchronization engine first, the team demonstrated the product as if it already worked and measured how many people asked to be notified — buying a read on desirability before writing the risky code. Simulation experiments sit between the two phases and are worth knowing by name: a Wizard of Oz test puts a real front end over a human-powered back end, and a concierge MVP delivers the value entirely by hand, so you learn whether people want the outcome before you automate a line of it.
The two phases compare like this:
| Dimension | Discovery experiments | Validation experiments |
|---|---|---|
| Purpose | Find direction, generate hypotheses | Confirm direction with real commitment |
| Cost & time | Low, fast | Higher, slower |
| Evidence strength | Weaker, more indicative | Stronger, more conclusive |
| Typical tools | Interviews, ads, fake doors, explainer videos | Single-feature MVP, concierge, pre-sale, split test |
| When to use | You know little; uncertainty is high | You've narrowed it; a real bet is next |
The takeaway: don't reach for an expensive validation experiment to answer a discovery-stage question — it's slow, and it presumes a direction you haven't earned yet. Climb the library from cheap and weak toward expensive and strong, and let each result decide whether the next, costlier rung is worth it. When you're unsure how much to trust a given result, the strength of your validation evidence is the deciding factor, not the effort you put in.
Sequencing Experiments From Weak to Strong Evidence
Sequencing is the skill that separates real validation from theater: you deliberately move from weak, cheap signals to strong, expensive ones, and you never treat a weak signal as a verdict. Bland and Osterwalder are blunt that not all evidence is equal — a survey answer and a signed pre-order are not the same currency, even when both say "yes."
The reliable dividing line is what people say versus what people do. Opinions, clicks, and sign-ups are easy to give and easy to walk back. Money, signatures, and repeated usage cost something, which is exactly what makes them trustworthy.
The evidence ladder sorts roughly like this:
| Signal | Weaker evidence | Stronger evidence |
|---|---|---|
| What it captures | What people say (opinions) | What people do (behavior) |
| Context | Hypothetical or artificial | Real and in-context |
| Commitment | A click, a like, an email | Money, a signature, sustained use |
| Sample | Small or self-selected | Larger or more representative |
Walk the clinic idea up that ladder to see sequencing in practice. You might start with interviews to confirm owners feel the no-show pain (cheap, weak-to-moderate). Then run a landing page measuring trial intent (moderate). Only then offer a warm segment an annual pre-pay at a discount (strong, because it's money on the table). Each rung is worth climbing only because the one below it held. If the interviews say nobody cares, you never spend on the pre-sale — the sequence saves you the expensive experiment precisely when the idea doesn't deserve it. How much a converting pre-sale should actually move your confidence is its own judgment call, and the strength of validation evidence is the guide to weighting each rung honestly.
The takeaway: stack experiments so that each strengthens the last. A landing page that converts (weak-to-moderate) earns a pre-sale test (strong); a pre-sale that lands earns a build. Combining and sequencing experiments this way — several per assumption, escalating in strength — is how you reach a confident decision without over-spending on any single test. One strong experiment can outweigh ten weak ones, but you usually can't afford the strong one until the weak ones have pointed the way.
Common Experiment-Design Mistakes
The most common failure isn't a bad experiment — it's a well-run experiment that couldn't have told you anything, because it was set up to confirm rather than to challenge. These are the recurring traps that turn testing into expensive reassurance.
- No pre-committed threshold. Without a criteria field filled in beforehand, every result looks "promising." Decide what pass and fail mean before you launch, or the data will mean whatever you need it to mean.
- Testing the comfortable assumption. Founders gravitate to easy, well-evidenced assumptions. If the test couldn't plausibly kill the idea, it isn't pointed at a real risk.
- Mistaking discovery for validation. Interviews and clicks are direction, not proof. Reading weak signals as strong is how teams scale a business nobody will pay for.
- Leading the witness. Asking "would you use a tool that does X?" invites a polite yes. Study behavior and past actions, not hypothetical future intentions.
- One test, one answer. A single experiment rarely settles an important assumption. Sequence a few, escalating evidence strength, before you commit.
- Skipping the decision. Capturing a learning and then not acting on it wastes the whole loop. Every learning card should end in "therefore, we will."
The through-line: design each experiment so a truthful "no" is genuinely possible and genuinely visible. If you can't describe the result that would make you stop, revisit which assumption you're testing and go back through your prioritized list in the complete guide to startup idea validation before spending another day running tests.
Key Takeaways
- Test assumptions, not the idea. Testing Business Ideas reframes validation as isolating the guesses a business depends on and buying evidence about the risky ones, rather than defending the idea as a whole.
- Prioritize with the assumptions map. Plot each assumption by importance and evidence; the high-importance, low-evidence quadrant holds your leap-of-faith assumptions and goes to the front of the queue.
- Cover all three risks. Desirability, feasibility, and viability each have to clear — and most teams over-test the feasibility they enjoy while under-testing the desirability and viability that actually kill companies.
- Commit the pass/fail bar in advance. The test card's criteria field, written before you run, is the strongest defense against rationalizing whatever result you happen to get.
- Close the loop with a learning card. Convert every observation into an explicit insight and a "therefore, we will" decision, or the experiment was activity, not learning.
- Match the experiment to the evidence you need. The library's 44 experiments span cheap-and-weak discovery tools to expensive-and-strong validation tools; start low and climb only as results justify it.
- What people do beats what people say. Money, signatures, and repeated usage are stronger currency than clicks or opinions, so sequence weak signals toward strong ones before you bet.
Frequently Asked Questions
What is Testing Business Ideas by Bland and Osterwalder about?
It is a practical field guide to de-risking new ideas through experiments rather than business plans. Published in 2019 as part of the Strategyzer series, it teaches you to extract your assumptions, prioritize them by importance and evidence, and run cheap tests that produce evidence about the risky ones before you invest heavily.
What are the 44 experiments in Testing Business Ideas?
They are a catalog of testing techniques the book rates by cost, evidence strength, and setup and run time, split into discovery and validation phases. Discovery experiments like interviews, fake-door tests, and explainer videos are cheap and fast; validation experiments like concierge MVPs, letters of intent, and pre-sales cost more but give stronger proof.
How do you decide which assumption to test first?
Test the assumption that is both most important to the idea working and least supported by evidence. Map every assumption on those two axes; the high-importance, low-evidence quadrant holds your leap-of-faith assumptions. Those go first, because that is where uncertainty is highest and being wrong is most expensive.
What is the difference between a test card and a learning card?
A test card is filled out before an experiment to define the hypothesis, the test, the metric, and the pass/fail criteria in advance. A learning card is filled out after, capturing what you observed, what you learned, and what you'll now do. One sets the bar; the other turns the result into a decision.
Is Testing Business Ideas worth it for solo and technical founders?
Yes — especially for technical founders prone to building before validating. The experiment library gives concrete, low-cost tests you can run without writing production code, and the assumptions map keeps you from wasting sprints on comfortable questions. It pairs naturally with lean-startup practice and works whether you're a solo founder or a corporate innovation team.