ICE Scoring: How to Rank Ideas by Impact, Confidence, Ease
ICE scoring ranks ideas by rating three factors — Impact, Confidence, and Ease — on a shared scale, usually 1 to 10, then combining them into one number. The bigger the score, the sooner you should run it. It trades precision for speed, which is exactly what an early backlog needs.
Quick Answer: ICE scoring is a lightweight prioritization method popularized by Sean Ellis for ranking growth experiments. You rate each idea on Impact (how much it moves the needle), Confidence (how sure you are), and Ease (how cheap it is to try), each 1–10, then compute ICE Score = Impact × Confidence × Ease. Highest score goes first.
If you have a list of twenty things you could build or test and no honest way to order them, you don't have a strategy — you have a mood. Whatever excited you last night wins. ICE exists to break that tie in about ninety seconds per idea.
It was popularized by growth-hacking pioneer Sean Ellis as a fast filter for a pipeline of growth experiments. The whole point is disposable speed: score an idea, rank it, run the top one, throw the number away. This guide covers what each factor means, how to combine them, when ICE is too blunt and you should reach for something heavier, and the mistakes that quietly turn the whole exercise into theater.
The Three ICE Factors: Impact, Confidence, and Ease at a Glance
ICE breaks every idea into three questions you can answer from memory. Impact asks how big the upside is if it works. Confidence asks how sure you are it will work. Ease asks how little it costs to find out. Each is a gut-calibrated number from 1 to 10, and the discipline is in answering all three honestly rather than any single one perfectly.
Here's the whole framework in one view before we take each factor apart. This table is a map of the questions, not a scorecard — the numbers come later.
| Factor | The question it answers | Score high (8–10) when… | Score low (1–3) when… |
|---|---|---|---|
| Impact | If this works, how much does it move my key metric? | It could meaningfully shift activation, revenue, or retention | Best case is a rounding error |
| Confidence | How sure am I this will actually work? | You have data, a prior win, or strong customer signal | It's a hunch with nothing behind it |
| Ease | How cheap and fast is it to try? | You can ship or test it this week, solo | It needs weeks of build or outside help |
The takeaway: a great ICE idea is one that's big, believable, and cheap all at once. Ideas that are huge but pure speculation, or safe but trivial, get filtered down automatically — which is the entire job. If you're using ICE upstream of building anything, pair it with the complete guide to startup idea validation so your Confidence scores rest on evidence instead of optimism.
Impact: How Much the Idea Moves Your Key Metric
Impact estimates the size of the upside if the idea works exactly as you hope. Rate it high when a win would visibly move the one metric you care about right now — signups, activation, revenue, retention — and low when even a total success barely registers. The trick is to always score against the same North Star metric so ideas stay comparable.
Impact is where founders lie to themselves most, so anchor it. Before scoring anything, name the single metric this quarter lives or dies on. Then ask each idea: in the best realistic case, how far does this push that number?
A rough calibration that keeps you honest:
- 9–10: Could plausibly change the trajectory of the whole metric (a new activation flow, a pricing change).
- 5–7: A solid, measurable lift, but not a step-change (a better onboarding email).
- 1–4: A nice-to-have, a polish item, or a bet on a metric that isn't your focus.
Two guardrails. First, score the realistic best case, not the fantasy — "this could 10x us" is almost never a 10. Second, don't let effort leak into Impact. A hard idea and an easy idea can have identical impact; that's what the Ease factor is for. Keeping the factors clean is what makes the final multiplication mean anything.
Confidence: How Sure You Are the Idea Will Work
Confidence rates how much evidence backs the idea, not how much you like it. Score it high when you have data, a past result, or a clear customer signal pointing the right way, and low when it's a hunch. This factor is ICE's built-in bullshit detector — it's the one that stops a wildly speculative moonshot from topping your list on Impact alone.
Be ruthless here, because Confidence is the factor people inflate to justify the idea they'd already decided to build. A useful ladder:
- 8–10: Direct evidence — an experiment, real usage data, or repeated unprompted customer requests.
- 4–7: Indirect signal — a competitor does it, a few interviews hinted at it, an analogy to a past win.
- 1–3: "I just think this'll work." No evidence, only conviction.
The honest question to ask out loud is: if I'm wrong about this, what told me? If the answer is "nothing, it just feels right," you're at a 2, not a 7. Scoring ICE confidence honestly is the difference between a ranking that protects you and one that just launders your existing bias into a number. When confidence is genuinely low but impact is genuinely high, that's not a reason to skip the idea — it's a signal to design a cheap test that raises the confidence before you commit real resources.
Ease: How Cheap and Fast the Idea Is to Try
Ease measures how little time, money, and effort the idea takes to ship or test — it's effort inverted, so easy scores high. Rate a one-afternoon experiment a 9 and a multi-week build with dependencies a 2. Framing effort as "ease" keeps the math consistent: for every factor in ICE, higher is always better, so you never have to remember which direction points to "do this."
That inversion matters and trips people up. In ICE, more effort means a lower score, because a big number in every column should always mean "go." Don't accidentally rate a huge, expensive build a 9 because it feels important — importance is Impact's job.
Calibrate ease against your actual capacity:
- 8–10: You (or one person) can ship or test it within a few days, no new dependencies.
- 4–7: A week or two of focused work, maybe one collaborator.
- 1–3: Multi-week effort, new infrastructure, or blocked on someone else.
For a solo founder or a tiny team, Ease often does the most sorting of any factor, because your real constraint is bandwidth. A "score" that's the same magnitude across ideas will get run first when it's a two-hour landing-page test versus a two-week feature — and it should. That bias toward cheap, fast tests is a feature of ICE, not a bug.
How to Score and Rank a Backlog With ICE
To rank a backlog with ICE, list every idea, score each on Impact, Confidence, and Ease from 1 to 10, multiply the three numbers for an ICE score, and sort descending. The top of the list is what you do next. The entire loop should take minutes, not a planning offsite — speed is the point.
The standard formula is a straight product:
ICE Score = Impact × Confidence × Ease
With a 1–10 scale, scores land between 1 and 1,000. Multiplication is deliberately punishing: a single low factor tanks the whole idea. An idea scored Impact 9, Confidence 9, Ease 1 lands at 81 — below a boring, safe, easy 7 × 7 × 7 = 343. That's ICE telling you a giant idea you can't cheaply de-risk should wait behind three things you can actually ship. (Some teams average the three factors instead of multiplying; averaging is gentler and won't let one weak factor sink an idea, but it also hides exactly the risk you want surfaced. Multiplication is the more common and more honest default.)
Let's rank a small hypothetical backlog. These numbers are illustrative — invented to show the mechanics, not benchmarks to copy.
| Idea | Impact | Confidence | Ease | ICE Score | Rank |
|---|---|---|---|---|---|
| Add a pricing page A/B test | 7 | 8 | 9 | 504 | 1 |
| Rewrite onboarding email sequence | 6 | 7 | 8 | 336 | 2 |
| Launch a referral program | 8 | 5 | 4 | 160 | 3 |
| Build a native mobile app | 9 | 6 | 1 | 54 | 4 |
The takeaway: the mobile app has the highest Impact of anything on the list, yet ranks dead last, because its Ease of 1 says you can't cheaply learn whether the impact is real. The pricing test wins not by being exciting but by being believable and nearly free to run. That's the ICE worldview in one table — reward what you can prove quickly.
A few working rules that keep the exercise fast and honest:
- Score all three factors before looking at the total. If you compute as you go, you'll fudge the last number to get the ranking you wanted.
- Do it in one sitting, with the same person or group calibrating. ICE scores are only comparable relative to each other; a 7 must mean the same thing across every row.
- Re-score, don't defend. When new evidence arrives, change the number. The score is a snapshot, not a contract.
For the full arithmetic and edge cases, see how to calculate an ICE score with worked examples.
When ICE Is Too Crude and You Need RICE
ICE is too crude when ideas reach wildly different numbers of people, because ICE has no factor for reach — a feature touching 1% of users and one touching everyone can score identically. When audience size varies a lot, upgrade to RICE, which adds Reach as an explicit fourth factor and grounds the estimate in numbers instead of gut feel.
The two frameworks share DNA; RICE is essentially ICE with a reach term bolted on and effort measured in real person-time rather than a 1–10 vibe. The trade-off is exactly what you'd expect.
| Dimension | ICE | RICE |
|---|---|---|
| Factors | Impact, Confidence, Ease | Reach, Impact, Confidence, Effort |
| Speed | Seconds per idea | Minutes per idea |
| Rigor | Gut-calibrated, subjective | Partly grounded in real numbers |
| Best for | Early, low-stakes, high-volume triage | Mature backlogs, cross-team buy-in |
The takeaway: ICE optimizes for how fast you can decide; RICE optimizes for how defensible the decision is. Pre-product-market-fit, when you're running dozens of cheap experiments and nobody needs a paper trail, ICE's speed wins. Once you have real user segments, a roadmap other people depend on, or a boss who wants to see the math, the reach factor stops being optional. Most founders start with ICE and graduate to RICE as the cost of being wrong goes up. For a side-by-side breakdown, read RICE vs ICE scoring.
Don't over-correct, though. A heavier framework doesn't make a guess more true — it just dresses it up in more decimals. If your Confidence scores are pulled from thin air, RICE's extra rigor is precision around a fantasy. Fix the evidence first.
Common ICE Scoring Mistakes That Break the Ranking
The most common ICE mistake is scoring the idea you already want to build up, so the "objective" ranking just confirms a decision you'd already made. ICE only works if you commit to the number before you know which idea it favors. A handful of predictable failure modes turn the framework from a decision aid into decoration.
- Reverse-engineering the score. You know you want to build feature X, so X mysteriously scores 9/9/9. If your gut has already decided, ICE can't help you — it can only rubber-stamp you. Score blind, then look.
- Inflating Confidence to justify excitement. A thrilling idea with zero evidence is Confidence 2, full stop. Conviction is not evidence. This is the single most abused factor, which is why scoring confidence honestly deserves its own discipline.
- Letting effort leak into Impact. Big, hard ideas feel impactful, so people over-score their Impact to compensate for the low Ease. Keep the factors independent or the multiplication is meaningless.
- Comparing scores across different sessions or people. An ICE score is only meaningful relative to the other scores set in the same sitting with the same calibration. A 400 from your co-founder last month isn't comparable to your 400 today.
- Treating the score as a verdict, not a prompt. ICE tells you what to try next, not what's true. The output of the top-ranked idea is data that should re-score everything below it. A ranking you never revisit is a ranking that's already stale.
- Using ICE for high-stakes, irreversible bets. A fast, subjective gut-score is perfect for "which experiment this week." It's the wrong tool for "should we raise a round" or "do we pivot the company." Match the rigor of the tool to the cost of being wrong.
The meta-mistake behind all of these: forgetting that ICE is a tie-breaker for cheap, reversible decisions, not a truth machine. Used for what it's good at, it's the fastest way to stop arguing and start testing. Used as a source of certainty, it just launders bias into three digits.
Key Takeaways
- ICE scoring ranks ideas on Impact × Confidence × Ease, each rated 1–10, and was popularized by Sean Ellis for prioritizing growth experiments — its whole value is speed, not precision.
- Impact is the size of the upside on your one key metric, Confidence is how much evidence backs the idea, and Ease is how cheap and fast it is to try (effort inverted, so easy scores high).
- Multiplying the factors is punishing on purpose: one low score sinks the whole idea, which is how a giant-but-unprovable bet correctly ranks below three cheap, believable ones.
- Score all three factors before you compute the total, in one sitting with consistent calibration, so you can't reverse-engineer the ranking you already wanted.
- Confidence is the most-abused factor — conviction is not evidence; score a thrilling but unproven idea a 2, and design a cheap test to raise it rather than faking it.
- Reach for RICE when ideas touch very different audience sizes, because ICE has no reach factor; ICE optimizes for decision speed, RICE for a defensible paper trail.
- ICE is a tie-breaker for cheap, reversible experiments, not a verdict on high-stakes bets — re-score as evidence arrives and match the tool's rigor to the cost of being wrong.
Frequently Asked Questions
What is the ICE scoring formula?
The standard ICE formula is ICE Score = Impact × Confidence × Ease, with each factor rated on a shared scale, most commonly 1 to 10. Multiplying yields a score from 1 to 1,000, and you rank ideas from highest to lowest. Some teams average the three factors instead, which is gentler but hides risk; multiplication is the more common and more revealing default because a single weak factor visibly tanks the total.
Who invented ICE scoring?
ICE scoring was popularized by Sean Ellis, a growth-hacking pioneer, as a fast way to prioritize a pipeline of growth experiments. It wasn't meant as a rigorous financial model — it was designed as disposable, gut-calibrated triage so teams could stop debating and start testing. That origin explains its character: it trades precision for the ability to rank a long list of ideas in minutes.
What scale should I use for ICE scores?
A 1-to-10 scale per factor is the most common and gives enough resolution without false precision. The scale itself matters less than using it consistently: a 7 must mean the same thing across every idea and every scorer in the session, because ICE scores are only meaningful relative to each other. Some teams use 1–5 for speed; pick one, write a quick rubric for what the endpoints mean, and stick to it.
Is ICE or RICE better for a startup?
Neither is universally better — it depends on stakes and audience variance. ICE is better early: when you're running many cheap, reversible experiments and nobody needs a paper trail, its speed wins. RICE is better once ideas reach very different numbers of users or your roadmap needs cross-team buy-in, because it adds an explicit reach factor and grounds effort in real person-time. Most founders start with ICE and graduate to RICE as the cost of a wrong call rises.
Why does one low factor ruin an ICE score?
Because the formula multiplies rather than adds. An idea with Impact 9, Confidence 9, and Ease 1 scores just 81 — well below a modest 7 × 7 × 7 = 343. That's intentional: a huge idea you can't cheaply test is a worse next move than a smaller one you can ship this week. The multiplication forces you to reward ideas that are big, believable, and cheap all at once, not just impressive on one axis.