Experiment Card Template: Design Any Validation Test

An experiment card is a one-page template that forces you to write your hypothesis, test method, metric, and pass mark before you run a validation test. In Testing Business Ideas, Strategyzer splits the job into two cards: a Test Card you fill in before, and a Learning Card you fill in after.

Quick Answer: An experiment card captures four things before you test: the hypothesis ("we believe that…"), the test ("to verify that, we will…"), the metric ("and measure…"), and the success criteria ("we are right if…"). Write all four in one sitting. If you cannot finish the criteria line, the test is not ready to run.

Why writing the card first beats running the test from memory

Writing the card first beats improvising because it forces you to name the pass mark before you have any data to argue with. That single act — deciding in advance what "success" looks like — is the whole reason the template exists.

Run a test without a card and three things tend to go wrong. You measure whatever is easiest to collect instead of what would actually change your mind. You interpret ambiguous results in favor of the idea you already like. And when a steering committee asks "how do you know?", you have a story instead of a threshold.

For an intrapreneur, that last problem is the expensive one. You are not just validating an idea for yourself; you are spending someone else's budget and defending the spend to people who did not sit in the room. A filled experiment card turns "the interviews felt positive" into "we set a bar of 10 of 15 and hit 9, so we are not funding the build yet." That sentence survives a review. A feeling does not.

The card is also cheap insurance against the most common failure mode in validation: moving the goalposts. Once you have seen the data, your brain will happily invent a threshold that the data clears. Pre-committing to a criteria line closes that door. It is the same discipline that makes a good step-by-step framework for validating any startup idea work — you decide what evidence counts before you go looking for it.

None of this requires software or a workshop. It requires four sentences and the honesty to write the criteria line before you are attached to the answer.

The experiment card template: hypothesis, test, metric, and criteria

The core template is four fields, each written as a sentence stem you complete. Fill them top to bottom and you have designed a test you can actually defend.

Here is the template with the stem for each field and the decision it forces you to make:

FieldSentence stemWhat it forces you to decide
Hypothesis"We believe that…"The one assumption this test is about — stated so it can be proven wrong
Test"To verify that, we will…"The cheapest action that produces real evidence, not an opinion
Metric"And measure…"The single number or signal you will actually collect
Criteria"We are right if…"The pass mark, set before you look at any results

The takeaway from the table: each field constrains the next, so a vague hypothesis produces a vague metric and an untestable criteria. Fix the top and the bottom fixes itself.

Around those four fields, the Strategyzer Test Card adds a handful of planning attributes that keep a portfolio of tests honest:

Those attributes matter most when you are sequencing several cards at once, which is why they live on the card and not in your head. Write the four core sentences first; add the planning attributes once the test itself is sound.

Field by field: how to fill hypothesis, test, metric, and criteria

Fill each field by moving from a belief to a number, in order, without skipping. Below is how to write each line so the card holds up.

Hypothesis: state one falsifiable belief, not a feature

The hypothesis names a single assumption in a form that could be proven false. "Regional dispatchers will love our tool" cannot fail — it has no edge. "Regional dispatchers lose more than an hour a day manually checking which shipments are at risk" can fail, because you can go count the hours.

Keep it to one belief per card. If your sentence contains an "and," you probably have two hypotheses and need two cards. When you are stuck turning a fuzzy idea into a sharp claim, work from a pattern — the formats in this guide to writing a good startup hypothesis name the who, the what, and the number you expect, which is exactly what the metric and criteria lines will need.

Test: pick the cheapest method that produces real evidence

The test is the action you take to check the belief. The rule is to pick the least expensive method that still generates evidence rather than an opinion. A five-question survey is cheap but often produces opinions; a landing page that asks for a real email or a card-on-file produces a signal.

Match the method to the hypothesis. A belief about a problem calls for interviews or observation. A belief about demand calls for a landing page, a presale, or a letter of intent. Do not reach for an expensive test when a cheap one would move your confidence just as far.

Metric: choose one number that would change your mind

The metric is the single signal you will collect. The test is that it should be a number you would act on — if the metric came back bad, you would genuinely stop. If no result would change your plan, you have picked a vanity metric and the card is theater.

One metric per card. Founders love to list five; then they cherry-pick the flattering one after the fact. Naming a single measure now removes that temptation later.

Criteria: set the pass mark before you look

The criteria line is where the card earns its keep. Write the exact threshold that separates "the belief held" from "it did not" — and write it before a single data point arrives.

Be specific and be honest. "Most people liked it" is not a criteria; "at least 8 of the 15 dispatchers raise this unprompted" is. If you cannot bring yourself to commit to a number, that discomfort is information: either the metric is wrong or you do not yet know what success would look like. Either way, the card is telling you the test is not ready.

The paired learning card: observation, insight, and decision

The Learning Card is the second half of the pair — you fill it in after the test to record what happened and what you will do about it. Where the Test Card looks forward, the Learning Card looks back, and its four fields mirror the first card so the two read as a matched set.

The Learning Card captures: the hypothesis you tested ("we believed that…"), the observation ("we observed…"), the learnings and insights ("from that we learned that…"), and the decisions and actions ("therefore, we will…"). The observation is raw and factual; the learning is your interpretation; the decision is the commitment that keeps the work moving. Skipping the decision line is how teams run a dozen experiments and still cannot say what any of them changed — the practice of capturing results on a learning card exists precisely so insight does not evaporate the moment the test ends.

This is where the terminology gets muddy, so it is worth being precise. Many teams use a single merged "experiment card" that stacks hypothesis, test, and result on one sheet filled in two passes. That works, but it blurs the seam between what you predicted and what you found. Strategyzer deliberately keeps them apart:

CardWhen you fill itThe four prompts, in order
Test CardBefore the testHypothesis → Test → Metric → Criteria
Learning CardAfter the testHypothesis → Observation → Learnings & insights → Decisions & actions

The takeaway: the split is a feature, not bureaucracy. Filling the Test Card before you run forces a prediction; filling the Learning Card after forces an honest comparison against that prediction. A merged card lets you quietly rewrite the prediction to match the result, which is the exact bias the whole method is designed to defeat.

A filled experiment card example (B2B intrapreneur)

Here is what a completed pair looks like for a realistic corporate scenario. Imagine an intrapreneur at a mid-size freight company who believes dispatchers waste time hunting for shipments about to miss their service-level agreement (SLA), and wants to pitch an internal tool that flags them automatically. The numbers below are the team's own illustrative targets, not research findings.

The Test Card (written before):

FieldWhat the intrapreneur wrote
HypothesisWe believe regional dispatchers lose meaningful time each day manually checking which shipments are at risk of missing SLA.
TestTo verify that, we will run 15 problem interviews with dispatchers across three regions and ask each to walk through their last at-risk shipment.
MetricAnd measure how many dispatchers name manual SLA-risk checking as a top-three daily frustration, unprompted.
CriteriaWe are right if at least 10 of 15 raise it unprompted and agree to join a pilot.

Notice what the card does before anyone is interviewed: it commits the team to a method (interviews, not a survey), a single metric (unprompted mentions), and a hard bar (10 of 15). There is nowhere to hide once the data lands.

The Learning Card (written after):

FieldWhat the intrapreneur wrote
HypothesisWe believed dispatchers lose meaningful time manually checking SLA-risk shipments.
Observation9 of 15 raised SLA-risk checking unprompted — just under the bar. But 12 of 15 described a deeper problem: the ETA data feeding their dashboard is often wrong.
Learnings & insightsThe SLA-risk symptom is real but downstream. The root frustration is ETA accuracy; flagging at-risk shipments off bad data would not help.
Decisions & actionsDo not build the SLA-flag tool. Write a new card to test the ETA-accuracy hypothesis with the same dispatchers next week.

This is the card doing its best work. The result technically missed the pass mark, and because the mark was written down, the team did not talk itself into building anyway. More valuable, the observation field surfaced a stronger signal the team was not testing for — and the decision field turned that into the next experiment instead of a hallway conversation that gets forgotten by Friday.

How to sequence multiple experiment cards into a test plan

Sequence your cards by risk and cost: test the assumption that would kill the idea first, and reach for the cheapest test that can move your confidence. One card is a test; a stack of cards ordered well is a validation plan.

Three rules keep the stack productive:

  1. Riskiest assumption first. Rank your beliefs by "what has to be true for this to work, and what am I least sure of?" Put a card on the intersection. There is no point perfecting the pricing card if the problem card has not passed.
  2. Cheap and weak before expensive and strong. Early cards use interviews and landing pages to cheaply kill bad ideas. Later cards use presales and pilots to expensively confirm good ones. Spend evidence-gathering money in proportion to how much confidence you have already earned.
  3. Let each Learning Card write the next Test Card. As the example above showed, the decision field of one card should point directly at the hypothesis of the next. When it does, your experiments compound instead of scattering.

Keeping the stack somewhere visible — a wall, a shared doc, or a validation workspace like Edmired — matters more than the tool you choose. The point is that anyone who asks "what have you learned and what are you testing next?" can see the answer without a meeting.

Adapting the experiment card for interviews, landing pages, and presales

The four fields stay the same across test types; only what you put in the metric and criteria lines changes. The template is deliberately format-agnostic, which is why the same card works for a conversation and a checkout page.

This table shows how the fields flex across three common early tests. It carries no invented figures — the criteria column shows the shape of a good pass mark, which you fill with your own number.

Test typeHypothesis focuses onMetric you would captureCriteria phrasing
Customer interviewsWhether a problem is real and painfulShare of interviewees who raise the problem unprompted"We are right if at least [N] of [total] describe it as a top frustration"
Landing page / smoke testWhether there is demand for the promiseShare of visitors who take a real action (email, waitlist, click to buy)"We are right if conversion clears [X%] over [traffic volume]"
Presale / letter of intentWhether people will commit before it existsNumber of paid preorders or signed intents"We are right if at least [N] commit within [time window]"

The takeaway: as you move down the table, the evidence gets stronger and the test gets more expensive. Interviews reveal whether a problem exists; a presale reveals whether anyone will actually pay to solve it. Sequence your cards down this ladder, and note the evidence-strength attribute on each so a reviewer can see at a glance how hard each result was to earn.

A word of caution on interviews specifically: your metric should count behavior and unprompted signals, not agreement. "Would you use this?" invites politeness. "Walk me through the last time this happened" invites facts. Write the metric line so it can only be satisfied by something the person did or said without being led.

Key Takeaways

Frequently Asked Questions

What is the difference between a test card and a learning card?

A test card plans an experiment before you run it; a learning card records what the experiment taught you after. The test card holds your hypothesis, test method, metric, and success criteria. The learning card holds the same hypothesis plus your observation, insights, and the decision you made as a result. Used together, they document a single experiment from prediction to action.

Is an experiment card the same as a hypothesis?

No. A hypothesis is one field on the card — the belief you are testing. The experiment card wraps that hypothesis in three more fields (the test, the metric, and the pass-mark criteria) that turn a belief into something you can actually run and score. A hypothesis tells you what you think; the full card tells you how you will find out.

How many experiment cards do I need per idea?

Usually several, run in sequence rather than all at once. Start with a card on the riskiest assumption — often whether the problem is real — and let each result point to the next card. Most ideas need separate cards for the problem, the demand, and the willingness to pay, because a single test rarely covers all three at a useful strength of evidence.

What makes a good success criteria on an experiment card?

A good criteria is a specific threshold you commit to before seeing any data, phrased so a result either clears it or does not. "Most people liked it" fails that test; "at least 8 of 15 preorder within two weeks" passes. If you cannot commit to a number, the discomfort usually means your metric is vague or you have not decided what success actually looks like.

Where did the experiment card template come from?

The Test Card and Learning Card were popularized by Strategyzer — Alexander Osterwalder, David Bland, and colleagues — and are laid out in detail in Testing Business Ideas. The broader idea of designing a falsifiable test around a hypothesis traces back to the lean startup and customer development movements, but the specific four-field card format and the test/learning split are Strategyzer's contribution.