Most roadmap slides run the same chain, and the middle link is usually left blank.
A field session on this site already gave that missing link a name: the Phase 2 Problem. It's named after the South Park gnomes, whose own business plan is Phase 1: collect underpants, Phase 2: a question mark, Phase 3: profit. The “?” usually gets filled with whatever's easiest to point at, engagement, NPS, sign-ups, none of which are the actual outcome.
This report picks up where that session's five-step exercise leaves off. Instead of a generic chain, it maps one illustrative feature launch in full, stage by stage, naming the actual metric at every link and the proxy trap waiting at each one. The example runs the whole way through: a bank ships a “set a savings goal” feature.
The chain doesn't stop at “people tried it.” It ends at whether the goal is still there, still funded, months later, and that last link is the one most launch reviews never check.
Every link in that chain is a bet on a specific behavioural assumption. Any single bet failing can quietly kill the initiative before anyone notices.
Mapping the chain is how those assumptions get caught early, while they're still cheap to fix.
The story in four parts
A feature that shipped, a goal that was created, a screen that was viewed: all real, none of them proof the thing actually worked.
Aware isn't the same as able to find it. Tried once isn't the same as finished. Finished isn't the same as still there next month.
Sign-ups, opens, and starts are cheap to log. Whether a goal survives contact with real life takes longer to find out, which is exactly why it gets skipped.
A named metric at every link, checked before launch, catches the weak join before it costs a whole quarter's read on whether the feature worked.
The evidence at a glance
At Microsoft's own experimentation platform, roughly a third of features tested actually improved the metric they were built to improve.
Across List's applied research, most programmes that showed real promise in a small pilot lose most of that effect once rolled out for real.
Real bank data found daily-login engagement negatively correlated with financial wellbeing, the opposite of what the metric was assumed to mean.
A randomised field trial found simple reminder messages raised the amount people actually saved, months after the account was opened, not just at sign-up.
Program evaluators have a formal version of the gnomes' problem: an output is what an activity directly produces, a feature shipped, a training run, a goal created. An outcome is the actual change in behaviour or condition that output is meant to cause, further down the chain, and it's measured separately, not assumed. Counting outputs is easy. Confirming an outcome takes waiting, and checking.
This distinction has been formalised for decades in program logic models, most influentially by the W.K. Kellogg Foundation's own guide, which separates outputs from short-term, medium-term, and long-term outcomes explicitly, precisely so a program doesn't mistake “we ran the activity” for “the activity worked.” A shipped feature is an output. Whether anyone's actual saving behaviour changed because of it is the outcome, several links further down.
The rest of this report maps that gap for one real kind of feature, stage by stage, with a named metric and a named trap at each link. The example: a bank ships a “set a savings goal” feature, letting a customer name a goal, a target amount, and a date, with progress shown on the home screen.
The five stages below aren't five separate checks. They're one chain, and every link narrows the group that reaches the next one. The map applies that chain to the savings-goal example, starting from an illustrative cohort of 1,000 customers who actually noticed the feature. The numbers are illustrative, built to show the shape a funnel like this takes, not a measured result from any real product. Tap any stage to see the smaller steps hiding inside it. Even “send a notification” is really three separate things that all have to happen.
Notices the feature: sees the banner, the notification, or the homepage tile.
Opens the goal-setting screen unprompted, days after the campaign ended.
Names the goal, the first substantive input.
Finishes the rest of the setup: amount, date, funding source.
Still has the goal active and funded, 90 days later.
About 95 of the original 1,000 are still there three months later, roughly one in ten. Each link lost roughly half the group that reached it, the same range List found evaporating between a pilot and a real rollout. That's the number worth taking to a steering committee. The 1,000 at the top, and the raw count of goals ever created, both look better, and both tell you less.
The map above already broke this stage into three real steps: sent, delivered, opened. Before anyone can try a feature, they have to know it's there. The proxy trap here is counting impressions instead of people. A banner shown five times to the same customer who never noticed it logs as five impressions and zero real awareness. The metric that actually matters is the share of eligible customers reached at least once through a channel they'd plausibly notice, not the raw number of times the banner rendered.
Why can a banner get a million impressions and still fail at awareness?
Impressions count renders, not people. A push notification sent to 500,000 customers but opened by only 4,000 of them logs as 500,000 impressions, when the real count of people who actually became aware is closer to 4,000, a gap of more than a hundred times.
What it teaches: the metric to report is unique customers reached, once, through a channel with a measurable open or view action attached, not the count of times a message was technically delivered.
Only 62% of that group reached this stage, based on the map above: opens the app, finds the tile. Awareness fades. A customer who noticed the feature announcement in March may still want it in July, and by then it has to be findable on its own, not through the original banner. The proxy trap is measuring traffic to the feature without separating who arrived by being pushed there from who found it on their own. Only the second group tells you whether the feature is actually discoverable inside the normal app, once the launch campaign is over.
Why does a launch week's traffic number say nothing about month three?
Total visits to the goal-setup screen in the first week are almost entirely campaign-driven and tell you nothing about ongoing findability. The real metric is organic entry rate: the share of visits arriving through normal in-app navigation, measured after the launch campaign has stopped running.
What it teaches: a feature that only gets found while it's being actively pushed hasn't solved discovery. It's borrowed attention from the campaign, and that borrowed attention runs out.
The map showed 55% making it this far: screen opens, names the goal. This is the link where “clicked a button” most often gets mistaken for “used the feature.” The proxy trap is counting a tap on the entry point as trial. The person may have opened the flow and immediately backed out. A real trial means reaching at least the first meaningful step inside the flow, not just arriving at its front door.
Why doesn't a tapped button prove anyone actually tried the feature?
“Set a Goal” button taps include everyone who tapped out of curiosity and left on the very next screen. The real trial metric is the share who reached the first substantive input, naming the goal, not just the share who opened the entry screen.
What it teaches: the first screen of a flow is still marketing. Trial starts at the first screen that requires the customer to actually do something.
Another 53% dropped off before this stage, the map's own count: amount, date, funding source. A started flow and a completed one are different outcomes, and the gap between them is where most of a feature's real friction lives. The proxy trap is reporting “flow starts” as if it were adoption. The real metric is full completion: name, target amount, target date, and a funding source all set, in one sitting or across return visits.
Where does a savings-goal flow actually lose people?
A four-step flow (name, amount, date, funding source) reporting only an overall “started” count hides which specific step is losing people. Naming each step's own completion rate, not just the flow's aggregate, is what turns a launch review from a single disappointing number into an actual list of things to fix.
What it teaches: this is the field session's own Rube Goldberg check, applied literally: find the weakest individual join in the chain, not just the chain's overall pass rate.
The map's final link, after one more 53% drop: three checkpoints, 30, 60, and 90 days out. This is the link most launch reviews never reach, because it's the one that takes the longest to measure and the easiest to skip. A completed goal is not a kept one. A customer can finish the entire setup flow and quietly stop funding the goal, or delete it, within weeks, and a dashboard reporting “goals created” will never show that it happened.
Why does “goals created” say nothing about whether the feature is actually working?
“Goals created” is a one-time output that never decays, even after every single goal has been abandoned. The real metric is a cohort-based persistence rate: the share of goals created in a given month that are still active and have received at least one contribution 30, 60, and 90 days later.
What it teaches: persistence isn't automatic once onboarding completes, and it isn't free. A randomised field trial found that simple reminder messages measurably increased how much people actually saved months after opening an account, evidence that sustained saving behaviour responds to ongoing prompts, not just a one-time setup flow.
Everything above is really one recurring mistake, worn five different ways: picking a metric that's easy to collect and treating it as if it were the outcome that actually matters. That's the same failure this site's own report on experiment validity calls a construct-validity problem, a real number standing in for a concept it doesn't actually capture. “Engagement” standing in for “loyalty” and “goals created” standing in for “customers who are actually saving more” are the identical mistake, one in a research paper and one in a product dashboard.
This is the field session's own review checklist, formalised. Four checks, not forty, because a review step people will actually run under deadline has to survive being tired on a Friday.
Every stage above looks like one clean step. Underneath, each is a bet that a specific behavioural assumption holds.
A notification earns real attention. Intention survives without a cue. Today loses to a future goal. One more form field won't cost the customer. A completed setup stays funded on its own.
Mapping the chain like this names those assumptions and checks them before they become a quarter's wasted spend. A named assumption has a mechanism behind it, salience, a missing signpost, present bias, and a mechanism is something behavioural economics can be pointed at to fix, not just measure.