How to Choose a Metrics Framework for Validation
Choose a metrics framework by the question you need to answer this quarter, not by which one is popular. Use One Metric That Matters to focus a single validation stage, AARRR to find where a funnel leaks, Google's HEART framework to judge experience quality, and a North Star metric to align a growing team on one measure of value.
Quick Answer: Match the framework to the question. Focus a scrappy pre-PMF bet with One Metric That Matters, diagnose a leaking funnel with AARRR (Pirate Metrics), measure experience quality with HEART, and unite a team behind one value number with a North Star metric. The best framework is the one whose core question is the decision you owe next.
Why your metrics framework should follow the question, not the fashion
The right metrics framework is the one whose built-in question matches the decision you actually owe this week. These frameworks are not interchangeable dashboards. Each was designed to answer one specific kind of question, and adopting one out of fashion rather than fit produces numbers nobody acts on.
Every framework encodes a question underneath its acronym. AARRR asks where in the funnel are we losing people? HEART asks is the experience good enough to keep? A North Star metric asks are we delivering more of the core value over time? One Metric That Matters asks what is the single riskiest number for the stage we are in right now?
Read those four questions again and notice they are not the same question phrased four ways. They sit at different altitudes and belong to different moments in a company's life.
The common failure mode is adopting whichever framework a famous company blogged about, instrumenting everything it names, and then deciding nothing — because the framework was built to answer a question you were not actually asking. Measurement effort is not free. Every metric you commit to tracking is a small ongoing tax on your attention and your tooling.
After building a few companies, the pattern that repeats is this: the specific framework matters far less than the honesty of the question behind it. Pick the wrong one and you spend weeks instrumenting movement when your real doubt was quality, or auditing experience when your real doubt was whether anyone wants the thing at all.
That last case is worth naming, because metric choice sits downstream of a more basic question. If you have not yet confirmed the problem is real and someone will pay to solve it, no dashboard saves you — that is the work covered in the complete guide to startup idea validation, and it comes first. Frameworks measure a thing that exists; they do not conjure demand.
Business model shapes the answer as much as stage does. A transactional product with a clear checkout is a natural home for AARRR's funnel, because the stages map onto real steps a buyer takes. A workflow tool people live inside all day leans toward HEART, where task success and engagement carry more signal than a one-time conversion. A multi-sided marketplace or a broad platform, where many teams could optimize in conflicting directions, is where a North Star metric earns its cost. Let the model narrow the field before you compare features.
Before you choose, answer three questions honestly. What decision am I trying to make? What stage is the product actually at? And how many people need to align behind the answer? Those three answers narrow the field faster than any feature comparison.
AARRR, HEART, North Star, and OMTM compared at a glance
The four frameworks founders reach for most differ less in sophistication than in the question each was built to answer. The comparison below sets them side by side on origin, core question, shape, and best-fit moment — all qualitative distinctions, because these are design choices rather than benchmarks, and there are no scores to invent.
| Framework (originator) | Core question it answers | Shape of the model | Best-fit moment |
|---|---|---|---|
| AARRR / Pirate Metrics (Dave McClure) | Where in the funnel are we losing people? | Five-stage lifecycle funnel | A live funnel that underperforms and needs diagnosis |
| HEART (Google — Kerry Rodden and colleagues) | Is the user experience good enough to keep people? | Five quality dimensions plus a Goals-Signals-Metrics process | Refining a product people already use repeatedly |
| North Star Metric (Sean Ellis; codified by Amplitude) | Are we delivering more core value over time? | One headline value metric with input metrics beneath it | Aligning a growing team on a single direction |
| One Metric That Matters (Lean Analytics — Croll & Yoskovitz) | What is the single riskiest number right now? | One rotating focus metric tied to the current stage | The earliest, most uncertain validation work |
The table's real lesson is that these frameworks are complementary, not competing. They operate at different altitudes — a single focus metric, a durable direction metric, a diagnostic funnel, an experience audit — and they suit different stages. Reading straight down the "best-fit moment" column is the fastest way to shortlist the one or two worth instrumenting now, and to defer the rest until their question becomes your question.
When AARRR (Pirate Metrics) fits best
AARRR fits best once you have a live product with real traffic and need to find where users fall out. Dave McClure's Pirate Metrics breaks the customer lifecycle into five measurable stages — acquisition, activation, retention, revenue, and referral — so you can isolate the leakiest step instead of guessing at the whole.
The framework's core strength is diagnosis. When a funnel underperforms, AARRR tells you which joint is failing: are you struggling to get people in the door, to get them to first value, to keep them coming back, to convert use into money, or to turn happy users into a referral channel? Each stage is a separable hypothesis with its own fix.
That strength comes with a precondition. AARRR presupposes a funnel already exists. Before launch, with no traffic to segment, there is nothing to instrument, and forcing the framework early produces five empty buckets. It is a tool for optimizing something that runs, not for deciding whether to build it.
The best-fit moments are concrete: comparing acquisition channels, hunting an activation drop-off, or deciding which single stage to fix first when several look weak. The full stage-by-stage instrumentation — and the metrics worth tracking at each stage before product-market fit — is laid out in the AARRR metrics framework founder's guide.
Its blind spot is worth stating plainly. AARRR measures movement, not quality. A funnel can convert respectably while the underlying experience quietly disappoints, and pirate metrics alone will not surface that. Reach for AARRR when the honest question is "which step do we fix first?" — and reach for something else when the question is "is this actually good?"
When Google's HEART framework fits best
HEART fits best when the thing in doubt is experience quality, not funnel throughput. Developed at Google by Kerry Rodden and her colleagues, HEART measures five dimensions of user experience — happiness, engagement, adoption, retention, and task success — and pairs them with a structured process for turning fuzzy UX into numbers you can actually track.
HEART's advantage is that it catches problems a funnel hides. Users can move through your conversion steps while quietly struggling: high effort, frequent errors, low task success, thin satisfaction. Those are experience failures, and they predict churn long before the funnel numbers admit it. HEART gives that intuition a measurable shape.
The practical engine is the Goals-Signals-Metrics method. For each dimension you care about, you first state the goal, then identify observable signals that indicate progress toward it, then choose the specific metrics that quantify those signals. It is what keeps HEART from becoming five vague adjectives. How to pick signals and goals for each dimension is worked through in the HEART framework UX guide.
You rarely track all five dimensions at once, and you are not meant to. HEART is a menu, not a checklist — a content site might live on engagement and happiness, while a productivity tool leans on task success and adoption. Pick the dimensions your product's value actually depends on.
The best fit is a product where the experience is the product: tools, workflows, anything used deeply and repeatedly, and teams who are refining rather than launching. HEART and AARRR make a natural pair — one asks whether the experience is good, the other whether the funnel flows — and many mature teams run both.
When a North Star metric fits best
A North Star metric fits best when a growing team is drifting apart across local metrics and needs one shared definition of value. Popularized by Sean Ellis and later codified in Amplitude's North Star playbook, it names the single number that best captures the value customers actually get — the metric every team can watch its work roll up into.
Its purpose is alignment and direction, not diagnosis. A North Star answers are we delivering more of the core value over time? and gives a marketer, an engineer, and a support lead one number they are provably contributing to. That shared gravity is the whole point; it is what a pile of disconnected team dashboards cannot provide.
What makes it operational is the input-metrics tree beneath it. The North Star itself is a lagging outcome — you cannot move it directly. The inputs are the handful of leading levers you can move, and mapping them is what turns an inspirational number into a working model. Choosing a North Star that reflects real value, and building those inputs beneath it, is the subject of the North Star metric founder's guide.
The classic mistake is choosing a vanity metric — registered users, cumulative signups — instead of a value metric that only rises when customers succeed. A good North Star should be nearly impossible to grow without genuinely helping people.
The timing caveat matters most for early founders. A North Star adopted before you have real product-market-fit signal simply enshrines a guess, aligning the team behind a definition of value you have not yet confirmed. It is a framework for scaling a validated direction, not for finding one.
When One Metric That Matters (OMTM) fits best
One Metric That Matters fits best at the earliest, most uncertain stage, when focus is scarcer than data. Drawn from Alistair Croll and Benjamin Yoskovitz's Lean Analytics, OMTM holds that at any given moment there is a single metric you should care about more than any other — and, crucially, that it changes as your riskiest open question changes.
The strength of OMTM is forced focus. Pre-product-market-fit, tracking a broad dashboard is often a sophisticated way of avoiding a decision. Naming one metric — the number that would prove or kill your current riskiest assumption — concentrates a small team on the thing most likely to sink it.
OMTM is deliberately temporary, which is what most distinguishes it. The metric is stage-bound: you pick the one number that answers today's biggest doubt, drive it until the doubt resolves, and then rotate to whatever risk that resolution newly exposes. It is a discipline more than a dashboard. The idea, and how the metric shifts as you progress through the stages, is explained in the one metric that matters guide, part of the broader Lean Analytics measurement approach.
The best fit is the scrappiest context: a solo founder, a small team, an early experiment where instrumenting everything is neither possible nor useful. OMTM meets you where the data is thin and the temptation to measure broadly is a trap.
OMTM and the North Star metric are constantly confused, and the difference is stage and durability. OMTM is temporary and rotates with your current risk; a North Star is durable and points in one direction for years. Many teams run an OMTM early, then graduate to a North Star once the direction is validated — the two are a sequence as much as a choice.
How to combine frameworks without drowning in dashboards
Combine frameworks by layering them at different altitudes, not by tracking all of them at once. Pick one primary lens for your current stage, borrow from a second only where it answers a question the first genuinely cannot, and delete any metric no decision depends on. Sprawl is the enemy, not incompleteness.
The frameworks stack cleanly because they were built for different heights. A workable arrangement for many startups treats a North Star as the durable direction at the top, AARRR stages and HEART signals as candidate input metrics feeding it, and One Metric That Matters as simply whichever of those inputs is riskiest this month. Seen that way, you do not run four dashboards — you run one hierarchy with a rotating point of focus.
This is also the honest reading of how the frameworks relate. AARRR and HEART largely populate the middle layer, supplying the diagnostic funnel steps and experience signals; the North Star gives them a shared summit; OMTM keeps a small team from staring at all of it simultaneously. Nothing here forces you to choose one framework forever.
The discipline that makes any combination work is subtraction. Every metric on a screen is a claim on someone's attention, and attention is the scarcest resource an early team has. If you cannot name the decision a metric informs, the metric should come off the board — not for tidiness, but because an unread number is worse than no number, since it looks like coverage while providing none.
This is where a validation-first workflow earns its keep. Edmired can tie whichever metric you are currently betting on to the evidence behind it, so the number you watch reflects real signal rather than dashboard habit — the same honesty-of-inputs principle that makes any framework worth adopting.
The framework, in the end, is scaffolding for a decision, not the decision itself. Choose the one whose question is your question, layer in a second only when a second question is genuinely live, and keep the board short enough that every number on it changes what you do next.
Key Takeaways
- The question comes before the framework. Pick the model whose built-in question matches the decision you owe this quarter, not the one a famous company happened to blog about.
- AARRR (Dave McClure's Pirate Metrics) is a diagnostic funnel. It isolates which of acquisition, activation, retention, revenue, or referral is leaking, but it needs a live funnel and measures movement, not quality.
- HEART (Google's Kerry Rodden and colleagues) measures experience quality. Its five dimensions and Goals-Signals-Metrics process catch frustration and low task success that a funnel hides — best for products used deeply and repeatedly.
- A North Star metric (popularized by Sean Ellis, codified by Amplitude) aligns a team on value. One durable number with input metrics beneath it gives a scaling team shared direction, but adopting it before product-market fit just enshrines a guess.
- One Metric That Matters (Lean Analytics, Croll & Yoskovitz) forces early focus. It names the single riskiest number for your current stage and rotates as that risk resolves — ideal for solo founders and thin-data experiments.
- The four frameworks are complementary, not competing. They sit at different altitudes and stages, and a North Star can hold AARRR and HEART metrics as inputs while OMTM marks whichever input is riskiest now.
- Subtraction is the real discipline. Every displayed metric taxes attention, so any number that informs no decision should come off the board.
Frequently Asked Questions
What is the best metrics framework for an early-stage startup?
For the earliest stage, One Metric That Matters usually fits best, because it forces a small team to focus on the single riskiest number instead of a broad dashboard. As you gain product-market-fit signal and add people, a North Star metric becomes more useful for alignment, with AARRR and HEART supplying diagnostic detail beneath it.
Can you use AARRR and a North Star metric together?
Yes, and they pair naturally. The North Star metric sits at the top as your one measure of value and direction, while AARRR's five funnel stages serve as input metrics beneath it — diagnostics you open when the North Star moves and you need to know which lifecycle step is responsible. One gives direction; the other explains changes.
What is the difference between a North Star metric and One Metric That Matters?
The difference is durability and stage. One Metric That Matters is temporary and rotates as your riskiest question changes, so it fits early validation. A North Star metric is durable, pointing a team in one direction for years, so it fits scaling a validated product. Many teams run an OMTM first, then graduate to a North Star.
Which metrics framework did Google create?
Google created the HEART framework, developed by researcher Kerry Rodden and her colleagues to measure user experience at scale. It tracks five dimensions — happiness, engagement, adoption, retention, and task success — and uses a Goals-Signals-Metrics process to convert each into concrete, trackable metrics. It is distinct from AARRR, North Star, and OMTM, which came from other originators.
Do I need a metrics framework before product-market fit?
You need focus more than a formal framework before product-market fit. One Metric That Matters gives that focus without heavy instrumentation, concentrating you on the single assumption most likely to be wrong. Elaborate funnel or alignment frameworks are premature until you have a working product and enough traffic to make their stages meaningful.