Special Report

Your Roadmap Has a Phase 2. Most Skip Straight Past It.

Most roadmap slides run the same chain, and the middle link is usually left blank.

A roadmap chain: Customers, then Initiative, then a missing link marked with a question mark, then Profit.
Customers Initiative ? Profit

A field session on this site already gave that missing link a name: the Phase 2 Problem. It's named after the South Park gnomes, whose own business plan is Phase 1: collect underpants, Phase 2: a question mark, Phase 3: profit. The “?” usually gets filled with whatever's easiest to point at, engagement, NPS, sign-ups, none of which are the actual outcome.

This report picks up where that session's five-step exercise leaves off. Instead of a generic chain, it maps one illustrative feature launch in full, stage by stage, naming the actual metric at every link and the proxy trap waiting at each one. The example runs the whole way through: a bank ships a “set a savings goal” feature.

The chain doesn't stop at “people tried it.” It ends at whether the goal is still there, still funded, months later, and that last link is the one most launch reviews never check.

It looks like one step. It's actually five, and most of them are invisible.

Every link in that chain is a bet on a specific behavioural assumption. Any single bet failing can quietly kill the initiative before anyone notices.

Mapping the chain is how those assumptions get caught early, while they're still cheap to fix.

The story in four parts

A feature shipping is only the start

A feature that shipped, a goal that was created, a screen that was viewed: all real, none of them proof the thing actually worked.

Every link between them can quietly break

Aware isn't the same as able to find it. Tried once isn't the same as finished. Finished isn't the same as still there next month.

The vanity metric is usually the easiest one to collect

Sign-ups, opens, and starts are cheap to log. Whether a goal survives contact with real life takes longer to find out, which is exactly why it gets skipped.

Mapping the chain is the deliverable, not a formality

A named metric at every link, checked before launch, catches the weak join before it costs a whole quarter's read on whether the feature worked.

The evidence at a glance

Only about a third of tested ideas work

At Microsoft's own experimentation platform, roughly a third of features tested actually improved the metric they were built to improve.

50 to 90% of a pilot's effect evaporates at scale

Across List's applied research, most programmes that showed real promise in a small pilot lose most of that effect once rolled out for real.

The most “engaged” customers were the least financially well

Real bank data found daily-login engagement negatively correlated with financial wellbeing, the opposite of what the metric was assumed to mean.

A text reminder measurably increased real saving

A randomised field trial found simple reminder messages raised the amount people actually saved, months after the account was opened, not just at sign-up.

An output is not an outcome, and the gap between them has a name

Program evaluators have a formal version of the gnomes' problem: an output is what an activity directly produces, a feature shipped, a training run, a goal created. An outcome is the actual change in behaviour or condition that output is meant to cause, further down the chain, and it's measured separately, not assumed. Counting outputs is easy. Confirming an outcome takes waiting, and checking.

This distinction has been formalised for decades in program logic models, most influentially by the W.K. Kellogg Foundation's own guide, which separates outputs from short-term, medium-term, and long-term outcomes explicitly, precisely so a program doesn't mistake “we ran the activity” for “the activity worked.” A shipped feature is an output. Whether anyone's actual saving behaviour changed because of it is the outcome, several links further down.

Source: W.K. Kellogg Foundation (2004). Logic Model Development Guide. The same discipline this site's own Phase 2 Problem field session arrived at independently, from a whiteboard exercise rather than a formal guide.

The rest of this report maps that gap for one real kind of feature, stage by stage, with a named metric and a named trap at each link. The example: a bank ships a “set a savings goal” feature, letting a customer name a goal, a target amount, and a date, with progress shown on the home screen.

The same chain, mapped end to end

The five stages below aren't five separate checks. They're one chain, and every link narrows the group that reaches the next one. The map applies that chain to the savings-goal example, starting from an illustrative cohort of 1,000 customers who actually noticed the feature. The numbers are illustrative, built to show the shape a funnel like this takes, not a measured result from any real product. Tap any stage to see the smaller steps hiding inside it. Even “send a notification” is really three separate things that all have to happen.

Funnel diagram: an illustrative cohort of 1,000 customers narrows to 95 across five stages, awareness, discovery, trial, completion, and persistence, with the conversion rate and drop-off between each stage shown below. Each stage expands to show the smaller steps inside it.
1 Awareness 1,000customers

Notices the feature: sees the banner, the notification, or the homepage tile.

62% continue 380 don’t
2 Discovery 620customers

Opens the goal-setting screen unprompted, days after the campaign ended.

55% continue 280 don’t
3 Trial 340customers

Names the goal, the first substantive input.

53% continue 160 don’t
4 Completion 180customers

Finishes the rest of the setup: amount, date, funding source.

53% continue 85 don’t
5 Persistence 95customers

Still has the goal active and funded, 90 days later.

About 95 of the original 1,000 are still there three months later, roughly one in ten. Each link lost roughly half the group that reached it, the same range List found evaporating between a pilot and a real rollout. That's the number worth taking to a steering committee. The 1,000 at the top, and the raw count of goals ever created, both look better, and both tell you less.

Stage one: awareness, does anyone know it exists?

The map above already broke this stage into three real steps: sent, delivered, opened. Before anyone can try a feature, they have to know it's there. The proxy trap here is counting impressions instead of people. A banner shown five times to the same customer who never noticed it logs as five impressions and zero real awareness. The metric that actually matters is the share of eligible customers reached at least once through a channel they'd plausibly notice, not the raw number of times the banner rendered.

Stage 1

The real metric, and the fake one

Why can a banner get a million impressions and still fail at awareness?

Impressions count renders, not people. A push notification sent to 500,000 customers but opened by only 4,000 of them logs as 500,000 impressions, when the real count of people who actually became aware is closer to 4,000, a gap of more than a hundred times.

What it teaches: the metric to report is unique customers reached, once, through a channel with a measurable open or view action attached, not the count of times a message was technically delivered.

Stage two: discovery, can they find it when they actually mean to?

Only 62% of that group reached this stage, based on the map above: opens the app, finds the tile. Awareness fades. A customer who noticed the feature announcement in March may still want it in July, and by then it has to be findable on its own, not through the original banner. The proxy trap is measuring traffic to the feature without separating who arrived by being pushed there from who found it on their own. Only the second group tells you whether the feature is actually discoverable inside the normal app, once the launch campaign is over.

Stage 2

The real metric, and the fake one

Why does a launch week's traffic number say nothing about month three?

Total visits to the goal-setup screen in the first week are almost entirely campaign-driven and tell you nothing about ongoing findability. The real metric is organic entry rate: the share of visits arriving through normal in-app navigation, measured after the launch campaign has stopped running.

What it teaches: a feature that only gets found while it's being actively pushed hasn't solved discovery. It's borrowed attention from the campaign, and that borrowed attention runs out.

Stage three: trial, did they actually try it once?

The map showed 55% making it this far: screen opens, names the goal. This is the link where “clicked a button” most often gets mistaken for “used the feature.” The proxy trap is counting a tap on the entry point as trial. The person may have opened the flow and immediately backed out. A real trial means reaching at least the first meaningful step inside the flow, not just arriving at its front door.

Stage 3

The real metric, and the fake one

Why doesn't a tapped button prove anyone actually tried the feature?

“Set a Goal” button taps include everyone who tapped out of curiosity and left on the very next screen. The real trial metric is the share who reached the first substantive input, naming the goal, not just the share who opened the entry screen.

What it teaches: the first screen of a flow is still marketing. Trial starts at the first screen that requires the customer to actually do something.

Stage four: onboarding and completion, did they finish the actual setup?

Another 53% dropped off before this stage, the map's own count: amount, date, funding source. A started flow and a completed one are different outcomes, and the gap between them is where most of a feature's real friction lives. The proxy trap is reporting “flow starts” as if it were adoption. The real metric is full completion: name, target amount, target date, and a funding source all set, in one sitting or across return visits.

Stage 4

The real metric, and the fake one

Where does a savings-goal flow actually lose people?

A four-step flow (name, amount, date, funding source) reporting only an overall “started” count hides which specific step is losing people. Naming each step's own completion rate, not just the flow's aggregate, is what turns a launch review from a single disappointing number into an actual list of things to fix.

What it teaches: this is the field session's own Rube Goldberg check, applied literally: find the weakest individual join in the chain, not just the chain's overall pass rate.

The exercise itself: Finding the “Phase 2” every roadmap quietly skips has the full five-step whiteboard version of this check, for mapping any initiative, not just this one feature.

Stage five: persistence, does it survive contact with real life?

The map's final link, after one more 53% drop: three checkpoints, 30, 60, and 90 days out. This is the link most launch reviews never reach, because it's the one that takes the longest to measure and the easiest to skip. A completed goal is not a kept one. A customer can finish the entire setup flow and quietly stop funding the goal, or delete it, within weeks, and a dashboard reporting “goals created” will never show that it happened.

Stage 5

The real metric, and the fake one

Why does “goals created” say nothing about whether the feature is actually working?

“Goals created” is a one-time output that never decays, even after every single goal has been abandoned. The real metric is a cohort-based persistence rate: the share of goals created in a given month that are still active and have received at least one contribution 30, 60, and 90 days later.

What it teaches: persistence isn't automatic once onboarding completes, and it isn't free. A randomised field trial found that simple reminder messages measurably increased how much people actually saved months after opening an account, evidence that sustained saving behaviour responds to ongoing prompts, not just a one-time setup flow.

Karlan, D., McConnell, M., Mullainathan, S., & Zinman, J. (2016). “Getting to the Top of Mind: How Reminders Increase Saving.” Management Science, 62(12), 3393–3411

Choosing the wrong stage as “the” outcome is a construct-validity problem

Everything above is really one recurring mistake, worn five different ways: picking a metric that's easy to collect and treating it as if it were the outcome that actually matters. That's the same failure this site's own report on experiment validity calls a construct-validity problem, a real number standing in for a concept it doesn't actually capture. “Engagement” standing in for “loyalty” and “goals created” standing in for “customers who are actually saving more” are the identical mistake, one in a research paper and one in a product dashboard.

The research version of this exact failure: Construct validity: is the metric measuring the thing it claims to?, on Five Ways an Experiment Can Be Right and Still Wrong, including the real bank study behind the “engagement” example.

Before you ship a success metric, four checks

This is the field session's own review checklist, formalised. Four checks, not forty, because a review step people will actually run under deadline has to survive being tired on a Friday.

  1. Have you named the actual outcome leadership cares about, not the easiest one to report?“Engagement,” “sign-ups,” and “Project X” are never the real answer on their own.
  2. Have you mapped every link between the initiative and that outcome, out loud?Awareness, discovery, trial, completion, and persistence are five separate links, not one step called “adoption.”
  3. Does every link have its own real metric, not a feeling?“It's going well” isn't a metric. A named completion rate or persistence rate, per link, is.
  4. Have you checked whether the metric survives past the launch campaign?A number driven by a temporary push (a banner, a promotion) isn't evidence the underlying link actually works on its own.

Why this is worth the whiteboard time

Every stage above looks like one clean step. Underneath, each is a bet that a specific behavioural assumption holds.

A notification earns real attention. Intention survives without a cue. Today loses to a future goal. One more form field won't cost the customer. A completed setup stays funded on its own.

Mapping the chain like this names those assumptions and checks them before they become a quarter's wasted spend. A named assumption has a mechanism behind it, salience, a missing signpost, present bias, and a mechanism is something behavioural economics can be pointed at to fix, not just measure.