ICE Scoring for Experiments: Prioritize What to Test

ICE scoring prioritizes a backlog of validation experiments by rating each test on Impact, Confidence, and Ease from 1 to 10, then multiplying the three into a single score. You run the highest score first. It ports the growth-experiment triage popularized by Sean Ellis onto your riskiest questions, trading precision for speed.

Quick Answer: ICE scoring ranks validation experiments so you test your riskiest assumption first. Rate each experiment on Impact (how much a clear result changes your next decision), Confidence (how trustworthy the signal will be), and Ease (how cheap and fast it is to run), each 1–10, then compute ICE Score = Impact × Confidence × Ease. Highest score runs next.

If you are a PM who became a founder, you already know how to rank a backlog. The shock is that the backlog changed shape. A roadmap is a list of things to build; a pre-fit startup is a list of things you don't yet know — and the expensive mistake is executing confidently in the wrong direction.

ICE ports your prioritization reflex onto that new backlog, with experiments as the unit instead of features. This guide covers why experiments need ranking at all, what Impact, Confidence, and Ease mean once the thing you're scoring is a test, how to score each without flattering yourself, a worked example ranking five experiments, how ICE compares to RICE and PIE, and the mistakes that quietly turn the exercise into theater.


Why Founders Must Prioritize Experiments, Not Just Features

Founders need to prioritize experiments because time and runway are the binding constraints before product-market fit, and every week spent testing a low-risk assumption is a week you didn't spend de-risking the thing that can actually kill the company. Feature prioritization asks "what should we build next?" Experiment prioritization asks the prior question: "what do we most need to find out before we build anything?"

The two are easy to conflate because they share the same verbs — score, rank, sequence. But the unit is different. A feature backlog assumes the direction is right and argues over sequencing. An experiment backlog assumes the direction is a hypothesis and argues over which doubt to resolve first. Get that wrong and you can ship a flawless roadmap toward a product nobody wants.

The instinct to just start building is the real enemy. Every founder has more test-worthy assumptions than time to test them: does the problem hurt enough, will they switch, will they pay, can you reach them, can you actually deliver. You cannot chase all of them at once, and you shouldn't chase the loudest one — you should chase the one where a cheap, clear answer changes the most about your plan.

That is a prioritization problem, and ICE is a fast way to solve it. If you're still mapping which assumptions are worth testing in the first place, start with the complete guide to startup idea validation and bring the resulting list of doubts here to rank.


The Three ICE Factors for Experiments: Impact, Confidence, and Ease

For experiments, the three ICE factors read slightly differently than they do for features: Impact is the value of the information a test produces, Confidence is how trustworthy that information will be, and Ease is how little time and money it takes to run. Each is a 1–10 gut-calibrated number, and the discipline is answering all three honestly rather than any single one perfectly.

ICE was popularized by growth pioneer Sean Ellis to triage a pipeline of growth experiments, and the original factors describe the change you're testing: Impact is the upside on your North Star metric if the experiment wins, Confidence is how sure you are it will, Ease is how cheap it is to try. That framing works fine. But validation experiments exist to learn, not to win, so it pays to read each factor through the lens of the learning rather than the launch.

Here is the whole framework in one view, translated for a backlog of tests. Treat it as a map of the questions to ask, not a scorecard — the numbers come in the next section.

FactorThe question for an experimentScore high (8–10) when…Score low (1–3) when…
ImpactIf this test returns a clear result, how much does it change my next decision?A pass or fail would redirect the roadmap, the pitch, or the spendYou'd carry on the same either way
ConfidenceHow trustworthy and clear will the signal be?The test measures real behavior — clicks, sign-ups, moneyIt rests on what people say they'll do, or a tiny, biased sample
EaseHow little time, money, and setup does it take to run?You can launch and read it this week, soloIt needs weeks of build, real budget, or someone else's sign-off

The takeaway: the best experiment to run next is decisive, trustworthy, and cheap all at once. A test that would change everything but rests on a hunch, or one that's trivially easy but tells you nothing you'd act on, gets filtered down automatically. That filtering is the entire job.


How to Score Impact, Confidence, and Ease 1–10 Without Kidding Yourself

Score each factor on the same 1–10 scale using a written rubric you agree on before you look at any experiment, so the numbers mean the same thing across every row. The goal isn't a perfect estimate — it's an honest, consistent one. Below is a calibration ladder for each factor, plus the habits that stop subjectivity from turning into noise.

Scoring Impact: The Value of the Information, Not the Size of the Dream

Rate Impact by how much a clear result would change your next move, not by how exciting the underlying idea is.

The tell of a low-Impact experiment is that you can't name what you'd do differently for each outcome. If "it works" and "it flops" both lead to the same next step, the information is worthless no matter how fun the test is. Write the two branches down before you score — that habit alone kills most vanity experiments.

Scoring Confidence: Trust in the Signal, Not Faith in the Hypothesis

Rate Confidence by how trustworthy the test's signal will be, not by how much you believe the hypothesis. This is the factor founders most often get backwards when the thing being scored is an experiment.

Note the trap hiding here: if you're already 90% sure the hypothesis is true, that doesn't make Confidence a 9 — it makes the experiment's Impact low, because you're about to spend time confirming what you already know. Confidence is about the instrument, not the conclusion.

This is also where the evidence hierarchy earns its keep. A round of customer interviews and a small presale can probe the same demand question, but the presale's signal is far harder to fake, so it earns a higher Confidence even though the interviews feel more thorough. Politeness inflates words; it rarely inflates a payment.

Scoring Ease: Time, Money, and Dependencies to Run the Test

Rate Ease as effort inverted — the cheaper and faster the test, the higher the score — so that a big number in every column always means "go."

For a solo or two-person team, Ease quietly does most of the sorting, because bandwidth is the true constraint. That bias toward cheap, fast tests is deliberate — you learn more by running five scrappy experiments than by perfecting one, and ICE is built to reward exactly that.

ICE is subjective by design, which is also its main weakness. Three cheap habits keep the subjectivity from becoming noise:

  1. Agree the scale first. Write one line each for what a 1, a 5, and a 10 mean per factor before scoring anything.
  2. Score in one sitting with the same people. A 7 has to mean the same thing across every row and every scorer.
  3. Re-score after results, don't defend the old number. Every experiment you run should move the Confidence and Impact of the ones still in the queue.

Because scores are only as good as the tests behind them, spell out each experiment's hypothesis, metric, and pass mark on an experiment card before you score it — a vague test invites vague, flattering numbers.


ICE Worked Example: Ranking Five Validation Experiments

To rank experiments, score each on Impact, Confidence, and Ease, multiply for an ICE score between 1 and 1,000, and sort descending. The hypothetical early-stage backlog below shows how the multiplication reorders your instincts. These numbers are illustrative — invented to show the mechanics, not benchmarks to copy.

The formula is a straight product:

ICE Score = Impact × Confidence × Ease

Validation experimentImpactConfidenceEaseICE ScoreRank
Fake-door button for a new feature on live traffic6894321
Landing-page smoke test with a small ad spend7783922
Pre-order page asking for real payment9842883
Five customer discovery interviews6582404
Build a functional MVP and watch retention9621085

The MVP build shares the highest Impact on the list and still ranks dead last, because an Ease of 2 says you can't cheaply learn whether that impact is real — the multiplication punishes it hard. The pre-order, which produces the most trustworthy signal money can buy, lands mid-table because it takes genuine setup. And the fake-door test wins not by being the most important but by being decisive enough, believable, and nearly free to run.

Notice the interviews. They're cheap and rank respectably, but their Confidence of 5 reflects that people are polite and stated intent is soft evidence — the same demand question, asked through a pre-order, scores higher precisely because dollars don't flatter you. That gap is the evidence hierarchy doing its job inside the score.

Run the top one, record what you learn, and re-score the rest. For a fuller version of this exercise with commentary on every factor, see the ICE scoring worked example for experiments.


ICE vs RICE vs PIE for Prioritizing Experiments

ICE, RICE, and PIE are all fast scoring models, and the right one depends on your stage: ICE for quick triage of a mixed validation backlog, RICE when experiments reach very different audience sizes or stakeholders want the math, and PIE when you're prioritizing A/B tests on live-traffic pages. All three combine a handful of 1–10 factors; they differ in what they force you to account for.

RICE, from Intercom, keeps Impact and Confidence, swaps Ease for Effort measured in real person-time, and adds Reach — the number of users a test or change touches — as an explicit factor. That matters because ICE is blind to reach: a test on a page 1% of users see and one everyone sees can score identically. PIE, from conversion-optimization firm WiderFunnel, scores Potential (how much room there is to improve), Importance (how valuable the traffic on that page is), and Ease (how hard it is to implement) — it's purpose-built for ranking optimization tests once you already have traffic.

The trade-off is speed versus what each model refuses to let you ignore.

ModelFactorsForces you to weighBest forOrigin
ICEImpact, Confidence, EaseNothing beyond gut-calibrated 1–10sFast triage of a mixed, early experiment backlogSean Ellis (growth experiments)
RICEReach, Impact, Confidence, EffortAudience size and real person-timeMature backlogs where reach varies or others need the mathIntercom
PIEPotential, Importance, EaseThe value of the page and traffic being testedPrioritizing A/B tests on live, high-traffic pagesWiderFunnel

For a pre-fit founder ranking a grab-bag of interviews, smoke tests, and presales, ICE's speed almost always wins — the tests are cheap and reversible, and nobody needs a paper trail. Reach for RICE once your tests touch very different numbers of users or a stakeholder wants defensible numbers, and for PIE only once you have live traffic and you're optimizing rather than validating. For the head-to-head that matters most early, read ICE vs RICE for experiments.

Don't mistake a heavier framework for a truer one, though. RICE's decimals and PIE's extra column don't make a guessed input correct — if your Confidence rests on nothing, more structure just adds precision to a fantasy. Fix the evidence before you upgrade the spreadsheet.


Building and Re-Scoring an Experiment Backlog With ICE

An experiment backlog is a running, ICE-ranked list of candidate tests — one per risky assumption — that you re-score every time a result lands. The score is a snapshot, not a contract: a single experiment's outcome should shift the Impact and Confidence of everything still in the queue, so a backlog you never revisit is one that's already stale.

The loop is small enough to run weekly:

  1. Capture. Every time an assumption surfaces — in a standup, an interview, a sales call — add it to the list as a candidate experiment, phrased as a test with a hypothesis.
  2. Score. Give each new candidate an Impact, Confidence, and Ease in the same sitting as the others, so the ranking stays comparable across old and new rows.
  3. Run the top one to three. Respect your Ease scores — pull the cheap, decisive tests first and keep a steady cadence rather than one heavy test a month.
  4. Re-score. When a result comes back, update the queue: a validated assumption often raises the Impact of the next test in the chain, while a failed one can delete three downstream experiments outright.

The board mechanics — columns, states, and cadence — deserve their own treatment; see how to build and manage an experiment backlog for the pipeline itself. ICE is just the ranking function that decides what rises to the top of it.

Wherever you keep the list — a spreadsheet, a Trello board, or a validation workspace like Edmired — the non-negotiable is that the scores stay visible and get re-scored on a cadence, not set once and forgotten. Re-scoring is the habit that separates ICE-as-decision-aid from ICE-as-decoration. Because the whole model is disposable by design, the number's job is done the moment it points you at the next test; then you throw it away and compute a fresh one against what you just learned.


Common ICE Scoring Mistakes That Distort Your Experiment Backlog

The most common ICE mistake is scoring the experiment you already want to run up, so the "objective" ranking just ratifies a decision you'd already made. ICE only works if you commit to the numbers before you see which test they favor. A few predictable failure modes turn the framework from a decision aid into theater — several of them specific to scoring tests rather than features.

The thread running through all of these: ICE is a tie-breaker for cheap, reversible experiments, not a truth machine. Used for what it's good at — ordering a long list of doubts you can test this quarter — it's the fastest way to stop arguing and start learning. Used as a source of certainty, it just launders your existing bias into three digits.


Key Takeaways


Frequently Asked Questions

How do you use ICE scoring for experiments?

List every candidate experiment, rate each on Impact, Confidence, and Ease from 1 to 10, multiply the three into an ICE score, and run the highest first. For tests, read Impact as how much a clear result changes your next decision and Confidence as how trustworthy the signal will be. Score all three factors before computing the total, so you can't reverse-engineer the ranking you already wanted.

What does Impact mean when scoring an experiment?

For an experiment, Impact is the value of the information — how much a clear pass or fail would change what you do next, not how exciting the underlying idea is. A test whose result you'd ignore either way scores low even if the feature behind it is huge. The quick check: write down what you'd do for each outcome; if the two branches match, Impact is near zero.

Is ICE or RICE better for prioritizing validation experiments?

ICE is usually better for early validation: when you're running many cheap, reversible tests and nobody needs a paper trail, its speed wins. RICE is better once experiments reach very different numbers of users or stakeholders want defensible numbers, because it adds an explicit Reach factor and measures effort in real person-time. Most founders start with ICE and graduate to RICE as the stakes rise.

Why does one low score sink an ICE experiment score?

Because the formula multiplies rather than adds. An experiment scored Impact 9, Confidence 9, Ease 1 lands at just 81 — below a modest 7 × 7 × 7 = 343. That's intentional: a high-impact test you can't run cheaply is a worse next move than a smaller one you can finish this week. The multiplication forces you to reward tests that are decisive, trustworthy, and cheap all at once.

How often should you re-score an experiment backlog?

Re-score whenever a result lands, and do a full pass at least weekly if you're running a steady cadence. Each completed experiment changes what you know, which changes the Impact and Confidence of the tests still queued. Because ICE scores are disposable snapshots, not commitments, updating them is the whole point — a backlog scored once and never revisited quickly ranks answered questions above open ones.