Willingness to Pay: How PMs Test Pricing Before Building
Test willingness to pay before building by shifting from what people say to what they'll actually do. Stated WTP — the answer to "would you pay $X?" — inflates, because agreeing costs the respondent nothing. Behavioral WTP, where someone parts with money, time, or reputation, doesn't lie the same way. Climb from pricing interviews toward pre-sales that quote a real number.
Quick Answer: Climb the willingness-to-pay ladder in order — pricing interviews to learn how buyers value the problem, structured surveys (Van Westendorp, Gabor-Granger) to frame a plausible range, then pre-sales that ask for money or a signed commitment. Each rung trades reach for realism; trust the rung where someone actually pays.
Why PM founders systematically under-test pricing
PM founders under-test pricing because their previous role almost never made them own the number. Inside an established company, price typically lived with finance, sales operations, or a pricing council, and often it was simply inherited from a model set years earlier. A product manager's mandate was activation, retention, and roadmap — not the figure on the invoice.
That division of labor leaves a specific blind spot. A PM-turned-founder arrives with excellent instincts for discovery, user problems, and iteration, then treats price as a launch-day setting rather than a hypothesis to validate alongside the product itself. The muscle for shipping features is strong; the muscle for extracting a real number from a real buyer is underdeveloped.
There's an emotional layer too. Asking someone to pay feels like selling, and selling is exactly the activity many PMs were content to hand to a go-to-market team. So pricing gets deferred — first we build, then we figure out what to charge — which is precisely backwards. If nobody will pay a viable number, that is the cheapest possible thing to learn before writing code.
The analyst's reframe is simple: price is a hypothesis with an evidence level attached. Testing Business Ideas by David Bland and Alexander Osterwalder is useful here because it ranks experiments by the strength of the evidence they produce — a call to buy beats a click, a click beats an opinion. Pricing deserves the same ladder.
Pricing is a validation input, not a launch decision
Pricing is a validation input because the willingness to pay a specific number is itself part of the problem you are testing. A feature nobody values at a sustainable price is not a pricing problem to solve later; it is a demand problem to discover now. Treating the number as testable — early, cheaply, repeatedly — is the whole shift a PM founder needs to make.
The four methods for testing willingness to pay before building
There are four practical methods for testing willingness to pay before you build, and they differ mainly in the kind of evidence they produce. Two are conversational or survey-based (stated preference), one is transactional (behavioral), and each has a place depending on how far along your thinking is. Choosing well among them is the core discipline of good willingness-to-pay research.
Here is how the four compare on what they ask, what they are best for, and how strong their evidence is.
| Method | What it asks the buyer | Best used when | Evidence type | Relative effort |
|---|---|---|---|---|
| Pricing interviews | How they value the problem, their budget, and current alternatives | You are still learning the buyer and the job-to-be-done | Stated, qualitative | Low to moderate |
| Van Westendorp PSM | Four price-perception questions (too cheap → too expensive) | You need a plausible range and an acceptable band | Stated, quantitative | Moderate |
| Gabor-Granger | Purchase likelihood at specific, escalating price points | You have a candidate price and want a demand-style curve | Stated intent, quantitative | Moderate |
| Pre-sales | For money, a deposit, or a signed commitment now | You have something concrete enough to sell | Behavioral | Moderate to high |
The takeaway: the first three methods tell you where to aim, while only the last tells you whether you will actually hit anything. Treat interviews and surveys as ways to earn a smart pre-sale, not as substitutes for one.
Pricing interviews: what they measure and when they mislead
Pricing interviews measure how a buyer thinks about the value, budget, and alternatives around a problem — not the exact number they will pay. Done well, they surface the mental model behind a purchase: who holds the budget, what the problem currently costs in workarounds, what "expensive" and "cheap" even mean in this category. That context is what makes every later number interpretable.
The method is qualitative, so the sample is small by design. Think a handful to a couple dozen conversations per segment, continued until you stop hearing new patterns. You are after saturation, not statistical power. Michele Hansen's Deploy Empathy is the practical reference for running these without contaminating the answers; the core skill is getting people to narrate past behavior instead of predicting future behavior.
Where interviews mislead is the moment you ask for a prediction. "Would you pay for this?" invites a polite yes that costs the respondent nothing, and a founder hungry for validation will over-weight it. Anchor on the past instead: what have they already bought to solve this, what did it cost, why did they switch or stay. A well-run session leans on a bank of sharp pricing interview questions that keep the conversation on spent money and real workflows rather than hypothetical enthusiasm.
Treat the output as directional. Interviews narrow the range and expose the value story; they do not confirm that anyone will transact. Their real job is to make your next, more expensive test smarter.
The Van Westendorp Price Sensitivity Meter, explained for founders
The Van Westendorp Price Sensitivity Meter is a four-question survey technique that maps the range of prices buyers perceive as acceptable, rather than a single "right" price. Each respondent answers four questions about the same product: at what price it would be so cheap they would doubt its quality, cheap enough to be a bargain, starting to feel expensive, and so expensive they would rule it out.
Plotting the cumulative distributions of those four answers produces intersection points conventionally read as boundaries — an acceptable price range, plus points often labeled the point of marginal cheapness and marginal expensiveness. The exact interpretive labels matter less to a founder than the shape: you learn where the perception of "too cheap" and "too expensive" begins to bite.
Because it depends on distributions, Van Westendorp needs more respondents than an interview round. As a planning rule of thumb, think dozens to low hundreds per segment for the curves to look stable, drawn from a sample that actually resembles your buyer rather than whoever happens to be in your network. Run it on the wrong crowd and you will get clean-looking curves that describe nobody.
Its honest limitation is that every answer is stated and hypothetical. The meter tells you what people say feels reasonable, not what they will pay when a card is required. Use it to set a defensible starting range and to spot segments with very different price perception — then go earn behavioral evidence.
Gabor-Granger: pinning down purchase intent at real price points
Gabor-Granger measures stated purchase likelihood at specific price points, letting you sketch a demand curve and a revenue-oriented number. Instead of open-ended perception questions, you show a respondent a concrete price and ask how likely they are to buy; then you vary the price up or down and watch how intent changes across the range.
Aggregated, those responses approximate how demand falls as price rises, which you can translate into a rough sense of a revenue-maximizing point. That is its advantage over Van Westendorp: it is oriented toward a decision — pick a number — rather than a perceived band.
Like Van Westendorp, it needs a reasonably sized, representative sample — again, think dozens to low hundreds per segment — and it inherits the same core weakness. Stated intent at a price is not a purchase at that price; "very likely" in a survey routinely evaporates at checkout. Sequential price questions can also anchor respondents, nudging answers toward whatever number you showed first.
For a founder, Gabor-Granger is most useful as the last stated-preference step before you charge real money. It converts a range into a specific candidate number that a pre-sale can then confirm or destroy.
Pre-sales: the only test where money actually changes hands
Pre-sales are the strongest willingness-to-pay test because they require the buyer to do something costly — pay, deposit, or sign — before the product fully exists. This is where stated preference becomes behavioral evidence. A landing page that quotes a real price and asks for a card, a founder-led pitch that ends in an invoice or a signed letter of intent, an annual prepay at a discount: each converts opinion into commitment.
What pre-sales measure is the only thing that ultimately matters — conversion at a number. You quote a specific price, you ask for the transaction, and you observe whether someone crosses the line. Even a soft commitment, like a deposit or a signed LOI with a start date, carries far more weight than a stack of survey "very likelys," because it costs the buyer money or reputation to give.
The sample can be small and still be decisive. A handful of real commitments from your target segment tells you more than a large survey, precisely because each data point is expensive to produce. Testing Business Ideas would classify this as high-strength evidence for exactly that reason: the buyer now has skin in the game.
The discipline is to make the offer real without over-building. Quote a genuine price, set genuine terms, and be ready to honor them. Tools like Edmired exist to help founders structure this ladder from conversation to commitment, but the mechanism itself is simple: ask for the money before you have spent months earning the right to. Once a number holds up in a pre-sale, translating it into tiers and packaging is the next problem — a separate craft covered in how to price a first SaaS product.
Stated vs behavioral WTP: which evidence to actually trust
Stated and behavioral willingness to pay differ in one decisive way: how easy the signal is to fake. Stated evidence — surveys, verbal agreement, "I'd definitely use this" — is cheap for the respondent to give, and therefore cheap to inflate. Behavioral evidence costs the buyer something real, which is exactly why it is trustworthy. Ranking your signals by that cost is the fastest way to stop fooling yourself.
The table below ranks common signals from weakest to strongest by how much they cost the person producing them.
| Signal | Evidence type | What it actually proves | How easily it's faked | Weight to give it |
|---|---|---|---|---|
| "Would you pay for this?" — yes | Stated | Politeness, or mild interest | Trivially | Very low |
| Survey price selection | Stated | Perceived acceptable range | Easily | Low |
| "I'd definitely buy that" (verbal) | Stated intent | Enthusiasm in the moment | Easily | Low |
| Joining a waitlist for pricing | Weak behavioral | Curiosity, low-cost interest | Somewhat | Low to moderate |
| Clicking a real "Buy" / paid-plan button (fake door) | Behavioral | Intent strong enough to act | Harder | Moderate |
| Deposit or pre-order | Behavioral | Money committed ahead of delivery | Very hard | High |
| Signed LOI or annual prepay | Behavioral | Budget and decision authority | Very hard | Very high |
The takeaway: weight each signal by what it cost the buyer to produce, and never let a pile of low-cost "yeses" outvote a single high-cost commitment. One deposit should move you more than a hundred survey checkboxes.
Common willingness-to-pay testing mistakes (and how to avoid them)
The most common WTP testing mistakes share one root: mistaking a cheap signal for a costly one. A PM founder under validation pressure is especially prone to collecting agreeable answers and reading them as demand. These are the traps that most often produce confident, wrong pricing.
- Asking "would you pay?" The hypothetical yes is free to give and nearly worthless. Replace it with questions about money already spent, or with an actual ask for a commitment.
- Testing on a biased sample. Friends, your professional network, and other founders will flatter your idea. Recruit people who match the buyer and have no reason to protect your feelings.
- Samples too tiny for the method. A dozen interviews is fine for qualitative saturation but far too thin for a Van Westendorp curve. Match sample size to the method's purpose.
- Pricing without a value proposition. A number asked in a vacuum measures nothing. A buyer can only price what they understand, so establish the value story before the number.
- Confusing intent with commitment. "Very likely to buy" is not a purchase. Treat every stated intent as a hypothesis that only a transaction can confirm.
- Leading and anchoring. Suggesting a price, or sequencing questions so an early number frames the rest, biases the answer toward what you hoped to hear. Let the buyer reveal the number where you can.
- One number for every segment. Different buyers value the same product very differently. Segment your evidence before you average away the signal.
Avoiding these is less about technique than temperament: keep asking whether a given signal cost the person anything to produce.
Key Takeaways
- Stated WTP inflates; behavioral WTP does not. The answer to "would you pay?" is free to give, so weight every signal by what it cost the buyer to produce.
- Price is a hypothesis, not a launch-day setting. Test the number alongside the product; a price no real buyer will pay is the cheapest possible thing to learn early.
- Climb the ladder in order. Pricing interviews frame value, Van Westendorp frames a range, Gabor-Granger picks a candidate number, and pre-sales confirm it with money.
- Match sample size to method. A handful of interviews reaches qualitative saturation; distribution-based survey methods need dozens to low hundreds of representative respondents.
- A single pre-sale outweighs a stack of surveys. One deposit or signed LOI is high-strength evidence precisely because it is expensive for the buyer to fake.
- Your sample is usually the weak point. Friends and your own network will mislead you; recruit buyers with no incentive to be kind.
- Segment before you average. One blended price hides that two buyer types value the product very differently — and the average may describe neither.
Frequently Asked Questions
How do I test pricing before launching a product?
Test pricing before launch by laddering evidence from talk to transaction. Start with pricing interviews to understand how buyers value the problem and what they already spend, use a Van Westendorp or Gabor-Granger survey to frame a plausible range, then run a pre-sale — a real price with a real ask for money or a signed commitment. The pre-sale is what actually validates the number.
Is a survey enough to validate willingness to pay?
No. A survey measures stated preference — what people say feels reasonable — which reliably overstates what they will pay once a card is required. Methods like Van Westendorp and Gabor-Granger are useful for framing a range and comparing segments, but they cannot confirm demand. Treat survey results as a hypothesis to be confirmed by a behavioral test, such as a pre-sale or a paid pilot.
How many people do I need to test willingness to pay?
It depends on the method's purpose, not a universal number. Qualitative pricing interviews become useful at a handful to a couple dozen per segment — enough to stop hearing new patterns. Distribution-based survey methods like Van Westendorp need a larger, representative sample, typically dozens to low hundreds per segment. Pre-sales can be decisive with only a few real commitments, because each one is costly to produce.
What's the difference between Van Westendorp and Gabor-Granger?
Van Westendorp maps a range of acceptable prices using four perception questions, so it answers "what band feels reasonable?" Gabor-Granger tests purchase likelihood at specific price points to sketch a demand curve, so it answers "which number should I pick?" Van Westendorp suits early framing; Gabor-Granger suits pressure-testing a candidate price. Both are stated-preference methods, so both still need behavioral confirmation.
Can I just ask customers what they would pay?
You can ask, but do not trust the answer as a number. Directly asking "what would you pay?" produces hypothetical figures shaped by politeness and guesswork, not budget reality. Instead, ask what they have already paid to solve the problem, what tools they currently buy, and how they justified those purchases. Then convert the strongest signals into an actual offer and see who commits.
How do pre-sales prove willingness to pay when the product isn't built?
Pre-sales prove willingness to pay because the buyer commits something costly — money, a deposit, or a signed letter of intent — before delivery. That cost is the proof: people rarely part with money or sign their name for something they do not genuinely want. Quote a real price, ask for a real commitment on real terms, and honor it. A few such commitments outweigh any volume of survey enthusiasm.