How Strong Is Your Validation Evidence? A Guide
Strong validation evidence records what people actually did when something real was at stake; weak evidence captures what they said they might do in a hypothetical. The strongest signals share three traits: real behavior over stated opinion, meaningful commitment over casual interest, and decision-driving metrics over feel-good vanity numbers.
Quick Answer: Grade every validation signal by three tests — did they do it or just say it, did they put skin in the game, and does the metric actually change your next decision? The more yeses, the more weight the evidence deserves. Praise, surveys, and sign-ups sit near the bottom; pre-orders, repeat usage, and paid behavior sit near the top.
Why Most Founders Trust Weak Validation Signals
Most founders trust weak signals because weak signals feel good, arrive early, and confirm what they already hoped was true. A stranger saying "I'd absolutely pay for that" lands as validation, when it's usually just politeness with no cost attached. The pull toward flattering evidence is the single biggest reason confident founders build products nobody buys.
Three forces make this trap so easy to fall into.
- Weak evidence is cheap to collect. You can run a survey or gather compliments in an afternoon; you cannot fake a customer handing over money.
- Weak evidence confirms the plan. Confirmation bias steers you toward signals that agree with your idea and away from the ones that would kill it.
- Weak evidence looks like progress. A growing waitlist or a spike in sign-ups produces a number that goes up, which feels like momentum even when none of it converts.
The cost of misreading these signals is rarely visible until late. A founder who reads a sign-up spike as demand invests months of build time against a number that was never going to convert, then blames execution, positioning, or timing when the launch stalls. The evidence was weak all along; only the interpretation was strong. Grading a signal up front is far cheaper than discovering its true strength after shipping.
The problem isn't that weak evidence is worthless — early curiosity is a real, if faint, signal. The problem is treating a faint signal as a green light. A founder who mistakes a hundred survey "yeses" for a hundred customers is optimizing for the wrong number, and the correction usually arrives after the product is already built.
The fix is not to distrust everything. It's to grade evidence deliberately, so you always know how much weight a given signal can bear. That grading discipline sits at the center of every serious validation method, and it's the through-line of the complete guide to startup idea validation — the rest of this article makes the grading criteria explicit.
The Validation Evidence-Strength Ladder, From Weakest to Strongest
Validation evidence sits on a ladder, and the rung a signal occupies depends on how much real behavior and real commitment it captured. The lower rungs record what people say in a low-stakes, hypothetical frame; the higher rungs record what they do when a decision costs them something. Knowing the rung tells you exactly how far the signal can carry your conclusion.
The ladder below ranks common validation signals from weakest to strongest. The strength column is deliberately qualitative — these are relative positions, not scores, and the point is the ordering, not any precise value.
| Validation signal | What it actually proves | Evidence strength |
|---|---|---|
| Unsolicited praise ("I'd totally use that") | The person is being encouraging in conversation | Weakest |
| Survey or interview answers about hypothetical intent | A stated preference inside a hypothetical frame | Weak |
| Email sign-up on a landing page | Mild curiosity, with almost nothing at stake | Weak to moderate |
| Meaningful time or effort invested (setup work, long trials, referrals) | Active interest and attention, though no money changed hands | Moderate |
| Pre-order, paid deposit, or signed letter of intent | Willingness to commit money or reputation before delivery | Strong |
| Repeat purchase or sustained real-world usage | Actual behavior with real stakes, sustained over time | Strongest |
The takeaway: two signals that sound similar can sit rungs apart. "Fifty people said they'd buy" and "fifty people put down a deposit" describe very different realities, and only the second earns a build decision. Use the ladder to locate any signal you've collected before you decide how much to lean on it. The next three sections unpack the axes that set a signal's height on this ladder, drawn from the lean canon and detailed further in the guide to Testing Business Ideas.
You also don't have to start at the top. Early on, weak signals are useful for choosing what to test next — a lively waitlist tells you where curiosity lives. The mistake is stopping there. Treat each lower rung as a prompt to design a higher-rung test: a waitlist becomes a pre-sale, a survey answer becomes a live checkout, an interview compliment becomes a request to commit. Climbing the ladder deliberately is how a validation program compounds real evidence instead of accumulating comfort.
Says Versus Does: Why Behavior Outranks Opinion
Evidence of what people do almost always outranks evidence of what they say, because saying costs nothing and doing costs something. This is the first and most important axis in the strength-of-evidence framework that David Bland and Alexander Osterwalder lay out in Testing Business Ideas: opinions are weak evidence, and observed facts about behavior are strong evidence. A customer's prediction about their future self is a guess; their action is data.
The gap between the two is not a small correction. People systematically over-report intentions that flatter their self-image — they'll pay more for sustainability, adopt the healthier option, switch to the better tool — and then, in the moment of decision, do the convenient thing instead. Your survey captures the aspiration; the market captures the behavior. When they conflict, the behavior is right.
This is why the format of a validation question matters as much as its content. Consider the difference:
- Says: "Would you use a tool that automated your weekly report?" — invites a hypothetical, agreeable answer.
- Does: "Here's the tool — connect your data and run this week's report." — forces an action with a real cost.
The second question can't be answered politely. Either they connect their data or they don't, and that choice is worth more than any number of enthusiastic maybes. Whenever you can convert a "would you" into a "will you, right now," you move the evidence up a rung.
A practical caution: interviews and surveys still have a place, but their job is to surface problems and language, not to confirm demand. Use what people say to learn what to test, then design an experiment that measures what they do. For the deeper mechanics of that translation — including how to structure demand tests around behavior — see the breakdown of say-versus-do evidence in validation, which extends this axis into a full testing method.
Skin in the Game: Why Commitment Outranks Casual Interest
The more a person risks to give you a signal, the more that signal is worth — commitment is the multiplier that separates real demand from idle interest. Testing Business Ideas frames this as the size-of-commitment axis: a small investment of time, money, or reputation makes for weak evidence, while a large investment makes for strong evidence. Alberto Savoia sharpens the same idea in The Right It with his Skin-in-the-Game Caliper — a deliberate gauge of how much a person actually put on the line behind their expressed interest.
The logic is simple. Anyone will click a "notify me" button; far fewer will pay a deposit; fewer still will pre-order and wait. Each escalation filters out the merely curious and leaves the genuinely committed, so the signal gets purer as the stakes rise. A hundred free sign-ups and ten paid pre-orders are not comparable quantities — the ten are worth more.
Skin in the game takes several forms, and a strong experiment asks for as much of it as the stage allows:
- Money — a deposit, a pre-order, an annual plan paid upfront. The most legible form of commitment.
- Time and effort — importing data, completing a real setup, sitting through a long onboarding. Costly attention is a genuine signal.
- Reputation — publicly recommending you, making an introduction, putting their name to a testimonial before the product exists.
- Switching — abandoning an incumbent tool or workflow to adopt your rough version, which carries the risk that it won't work.
Savoia's broader warning in The Right It is that founders spend too long in what he calls "Thoughtland" — reasoning about whether an idea will work from opinions and speculation — instead of collecting what he calls YODA, or "your own data," from real-world tests. The Skin-in-the-Game Caliper is how you weight that data once you have it: a yes backed by commitment counts; a yes backed by nothing barely registers. When you design a validation test, the design question is always, "what am I asking this person to risk?"
Actionable Versus Vanity: Why Decision-Driving Metrics Outrank Feel-Good Numbers
A metric is only strong evidence if it changes what you do next; a number you'd never act on is vanity, no matter how large or how fast it grows. This is the axis Alistair Croll and Benjamin Yoskovitz build Lean Analytics around: vanity metrics make you feel good but drive no decision, while actionable metrics tie to a specific choice and tell you what to change. Total sign-ups, raw pageviews, and cumulative downloads are the classic vanity numbers — they only ever go up, and they never tell you what to do.
The test they propose is disarmingly direct: for any number you're tracking, ask, "what will I do differently based on this?" If the honest answer is "nothing," you're looking at a vanity metric dressed up as validation. A total that can only rise gives you no decision to make; a rate or a ratio that can move in either direction does.
The contrast is easiest to see side by side:
- Vanity: total registered users — always climbing, hides churn, prompts no action.
- Actionable: activation rate of new users this week — can fall, exposes a broken onboarding, tells you where to fix.
- Vanity: cumulative downloads — a lifetime sum that never goes down.
- Actionable: percentage of downloads that reach the core action — reveals whether the product delivers, and points at the leak.
Vanity numbers are seductive precisely because they're easy to grow and hard to make fall. You can always add another channel and watch total sign-ups tick up, which feels like progress even as the fraction who ever reach real value quietly shrinks. Actionable metrics resist that comfort: because they can drop, they force you to confront the part of the funnel that isn't working. In validation, the number you're a little afraid to look at is usually the one telling the truth.
A strong validation metric, in the Lean Analytics sense, is comparative (you can measure it against a segment or a prior period), it's usually a rate or a ratio rather than a running total, and above all it changes your behavior. When you report an experiment result, lead with the number that would have made you kill the idea had it come in low — that's the actionable one. The vanity numbers can go in a footnote, if anywhere.
A Rubric to Score Any Experiment Result
Score any experiment result by running it through four questions drawn from the three axes above; the more it satisfies, the more weight the evidence deserves. This turns "is this good enough?" from a gut feeling into a repeatable check you can apply to a survey, a landing page, a pre-sale, or a pilot alike. Run every result through the same filter and you stop grading evidence by how much you like the answer.
Ask, in order:
- Say or do? Did the person take a real action, or express an opinion about a hypothetical? Actions outrank statements every time.
- Skin in the game? What did the signal cost them — money, meaningful time, reputation, a switch? The higher the cost, the stronger the evidence.
- Actionable or vanity? Would a low result have changed your decision? If no number could have stopped you, the metric proves nothing.
- Real or artificial context? Did the behavior happen in a realistic setting, or in a lab-like situation the person knew was a test? Testing Business Ideas counts real-world context as stronger than an artificial one, because people behave differently when they know they're being observed.
A signal that clears all four — a real action, at real cost, on a metric you'd act on, in a real setting — is about as strong as early evidence gets. A signal that clears none is a compliment. Most results land in between, and the rubric's value is telling you where, so you calibrate your confidence to match. A fifth question is worth adding once you can: how repeatable is it? One committed customer is an anecdote; the same behavior across a segment is a pattern, and Testing Business Ideas treats more independent data points as stronger evidence than fewer.
Applied honestly, this rubric will sometimes downgrade a result you were excited about — that's the point. It's the same discipline that structures a rigorous validation workflow end to end, and pairing it with the complete guide to startup idea validation gives you both the grading criteria and the experiments that generate gradeable evidence in the first place.
A Worked Example: Scoring a Landing-Page Result Against a Pre-Sale
Run two real results through the rubric and the gap becomes obvious. Suppose a landing page collected two hundred email sign-ups, and a follow-up page collected twelve paid deposits. The sign-ups are a say signal — a click — with near-zero skin in the game, tracked by a running total you'd never act on, in a context the visitor half-knew was a test; they clear none of the four questions. The twelve deposits are a do signal, real money at stake, a conversion rate you'd absolutely act on, in a live purchase flow; they clear all four. The two hundred is the bigger number and the weaker evidence. Sizing your confidence to the twelve, not the two hundred, is the entire discipline in miniature — and it's why a tool like Edmired organizes evidence by strength rather than by volume.
Common False-Positive Validation Signals to Distrust
A false-positive signal is one that feels like validation but disappears the moment you ask for commitment — and a handful of them account for most premature builds. Each one shares a tell: it's abundant, cheap for the other person to give, and it flatters the idea. Learning to recognize them on sight is half the battle, because they're specifically the signals a hopeful founder wants to over-read.
Watch for these in particular:
- Compliments in interviews. "That's a great idea" is social lubrication, not demand. The interview is useful for the problem you uncover, not the praise you collect.
- Survey intent with no cost. A high "yes, I'd buy" rate on a survey measures agreeableness inside a hypothetical, and rarely survives contact with a real price.
- Warm-intro enthusiasm. Friends, colleagues, and people who like you will support you regardless of the product. Their goodwill is real; their objectivity isn't.
- A large waitlist that never activates. Email sign-ups sit low on the ladder for a reason — the cost to join is a click, and a big list with no conversion is a vanity metric, not proof.
- Free usage with no willingness to pay. People will happily use anything free. Enthusiastic free users who vanish at the paywall have told you about the price, not the demand.
- Total-count metrics that only rise. Cumulative sign-ups, downloads, and pageviews are the vanity numbers Lean Analytics warns about — they hide churn and prompt no decision.
The unifying diagnosis for all six is low commitment: nobody in these scenarios risked anything, so nobody's signal earned much weight. The corrective is always the same move — raise the ask. Turn the survey into a pre-order, the waitlist into a deposit, the free trial into a paid one, and watch how many of the "yeses" survive. The ones that do are worth more than everything you filtered out.
Key Takeaways
- Grade evidence, don't just collect it. Every validation signal carries a different weight; the founder's job is to know which rung of the ladder a signal sits on before leaning on it, not to treat all "yeses" as equal.
- Does beats says. What people do when something is at stake outranks what they predict they'll do, because saying is free and doing is not — when behavior and stated intent conflict, believe the behavior.
- Skin in the game is the multiplier. The more a person risks — money, meaningful time, reputation, a switch — the more their signal is worth; Savoia's Skin-in-the-Game Caliper is how you weight any yes you receive.
- Actionable outranks vanity. A metric is only strong evidence if a low result would have changed your decision; totals that only climb, like cumulative sign-ups and downloads, prompt no action and prove little.
- Context strength counts too. Behavior observed in a realistic setting beats behavior in an obvious test, and repeated evidence across a segment beats a single committed anecdote.
- Most false positives share one flaw: no commitment. Compliments, survey intent, warm-intro enthusiasm, idle waitlists, and free usage all feel like validation precisely because they cost the giver nothing — raise the ask and most evaporate.
- When in doubt, escalate the request. Turning a "would you" into a "will you, right now" is the fastest way to convert weak evidence into strong evidence and to find out what your demand is actually made of.
Frequently Asked Questions
What Is the Difference Between Weak and Strong Validation Evidence?
Weak validation evidence captures what people say they might do in a hypothetical, low-stakes situation — opinions, survey answers, and compliments. Strong evidence captures what people actually do when a decision costs them money, time, or reputation, in a realistic setting. The Testing Business Ideas framework grades evidence along exactly these axes: says versus does, small versus large commitment, and artificial versus real context.
Is Survey Data Ever Strong Enough to Validate an Idea?
Survey data is rarely strong enough to validate demand on its own, because a survey measures stated intent at zero cost, and stated intent routinely overstates real behavior. Surveys are genuinely useful for surfacing problems, language, and segments to test — but treat them as a source of hypotheses, not confirmation. Before you build, convert the survey's "yes" into a behavioral test that asks the person to actually commit.
How Much Skin in the Game Counts as Strong Evidence?
The more the signal cost the person, the stronger it is — there's no fixed threshold, only a relative ladder. A paid pre-order or deposit outranks a free sign-up; importing real data or switching away from an incumbent tool outranks a quick click. Savoia's Skin-in-the-Game Caliper is a mindset, not a number: for any yes you receive, ask what the person actually risked to give it, and weight it accordingly.
What Are Examples of Vanity Metrics in Startup Validation?
Classic vanity metrics include total registered users, cumulative downloads, raw pageviews, and social media likes — numbers that only ever rise and hide the churn underneath. Lean Analytics diagnoses them with one question: "what would I do differently based on this number?" If the answer is nothing, it's vanity. Replace each with an actionable counterpart, such as new-user activation rate or the share of trials that convert to paid.
Can Strong Evidence Still Be Wrong?
Yes — strong evidence lowers your risk of a false positive, but it never eliminates it. A committed pilot customer can be unrepresentative, an early cohort can behave differently from the mainstream, and a real-world test can still be too small to generalize. That's why repeatability is part of the rubric: strong evidence gets stronger when the same behavior recurs across independent people and segments, not just once with an enthusiastic early adopter.