A number produced by comparing before a change to after it looks like measurement. Usually it isn't. It can't tell a real effect apart from a season, a competitor's bad month, a trend that was already climbing, or ordinary noise settling back towards average.
That gap doesn't scale down as the bet gets bigger, it scales up. A $10M feature, product launch, or marketing campaign judged by a pre-vs-post number is a $10M decision that was never actually measured, dressed as one that was.
Four real cases follow: two where a company assumed an effect and found none, one where a whole industry has almost never checked, and one where the actual arithmetic shows why a real effect can hide from a comparison this crude, no matter how carefully it's read.
The story in four parts
Without a real comparison group, there's no way to know what would have happened anyway. The number describes a coincidence in time, not a cause.
A real field experiment, not a guess: cutting a $100M-a-year budget line moved sales by an amount indistinguishable from zero.
Real observational estimates and real experiments on the identical campaigns disagreed by factors of three, ten, or more, in both directions.
The honest cost of skipping measurement isn't recklessness. It's not finding out, later, whether the money did anything.
The evidence at a glance
eBay's own randomised experiment found brand-keyword paid search added no measurable sales. Its CEO cut the budget by roughly $100 million a year.
One Facebook case study found an observational model estimating 1,306% lift, against a real randomised result of just 2.4%.
Applied to eBay's own data, the same style of regression estimated +4,100% ROI. The real experiment measured -63%.
Across 25 real retail experiments, the median sample needed to tell “worked” from “didn't” was 3.3 million exposed customers.
“Sales were up 12% after we launched it” sounds like a finding. It's actually four different things bundled into one number, and a real experiment is the only design that separates them.
Without a group that didn't get the change, there's no version of events to compare the real one against. The 12% could be the change, or it could be what would have happened anyway.
A season, a competitor's outage, a macro shift, an unrelated feature shipped the same week. Pre/post has no way to separate the change from whatever else changed at the same time.
A change launched right after an unusually bad stretch will often look like it worked, purely because the bad stretch was never going to stay that bad.
A team runs several pre/post reads on the same launch, across different metrics and windows, and the one that looks best becomes the story that survives the debrief.
eBay ran brand-keyword paid search ads for years: pay to appear at the top of a Google or Bing search for “eBay,” directly above the identical free organic listing one line down. Internal belief, based on ordinary before/after tracking, was that these ads drove real incremental sales. Economist Steve Tadelis, working inside eBay, ran a real randomised field experiment instead: turning brand-keyword ads off in some markets while leaving them on in others.
Sales in the markets with ads off were statistically indistinguishable from the markets with ads on. Nearly every click the ad had been buying would have arrived anyway, through the free organic listing sitting right underneath it.
Tadelis pushed the test further, pulling ads for generic product-search terms too. The average effect on sales was, again, indistinguishable from zero. eBay's CEO subsequently cut the company's paid-search budget by roughly $100 million a year.
Nobody at eBay was being careless before this test ran. The pre/post numbers genuinely looked good, because a customer who clicked a brand-keyword ad usually was already about to buy from eBay anyway. Only a real experiment, a market with the ads simply switched off, could show that the same customer would have arrived for free.
The companion episode to eBay's paid-search story widens the lens: businesses worldwide spend hundreds of billions of dollars a year on advertising, and the overwhelming majority of that spend has never been checked against a real experiment, a market or audience with the ad simply withheld. Most of it runs on exactly the pre/post logic this report is about: a campaign launches, a number moves, credit gets assigned.
This isn't a claim that advertising never works, eBay's own later experiments, and the case below, both found real effects elsewhere. It's a claim about the ratio: an enormous amount of spend is justified by a measurement method that structurally can't rule out the alternative explanations in the taxonomy above, and comparatively little of it is ever tested against a design that can.
The first two cases both involve a plain before/after comparison, the crudest version of the mistake. A regression on historical data isn't that same crude tool, it controls for known confounders and looks more rigorous. A large study directly comparing regression-based estimates against 15 real, randomised Facebook ad experiments found that isn't enough to trust it either.
Comparing what a careful observational model estimated against what the matched randomised experiment actually found, one case showed a real measured lift of 2.4%, against an observational estimate of 1,306%.
Applied to eBay's own non-brand search data, the same regression approach estimated an ROI of over 4,100%. The randomised experiment above found an ROI of roughly −63%. The two methods didn't just disagree on magnitude, they disagreed on the sign.
The paper's own conclusion is the one worth sitting with: the errors don't run in a single, correctable direction. A regression estimate can overstate a real effect by an order of magnitude or invent one that isn't there at all, so there's no fixed discount rate to apply to an observational number to make it trustworthy. The only way to know which side of that gap a given estimate sits on is to run the actual experiment.
Even setting confounding aside entirely, there's a second, purely statistical problem: individual purchase behaviour is noisy, far noisier than most advertising budgets are sized to detect. A study of 25 real field experiments run with major US retailers and brokerages, together representing millions of exposed customers, measured exactly how much data a genuine effect needs before it stops looking like noise.
That's the median sample size the 25 experiments needed to distinguish a campaign that broke even from one that did nothing at all, a distinction most companies believe they can eyeball from a pre/post chart.
The sample size needed to reliably detect a genuinely real but modest 5% improvement, an effect size plenty of internal launch reviews claim to have found with nothing like this much data.
The paper's median confidence interval on a campaign's real return on investment spans more than 100 percentage points wide, even inside these large, well-designed experiments. A pre/post comparison, run on a fraction of that data with no control group at all, isn't a smaller, cheaper version of this measurement. It's reporting a single point where the honest answer is a range wide enough to contain both “this worked well” and “this lost money.”
None of the four cases above argue that every feature, campaign, or product decision needs a randomised trial before it ships. That bar is genuinely too high for most decisions: too slow for a small fix, too expensive for a low-stakes call, sometimes structurally impossible to run cleanly at all. The argument is narrower and harder to talk your way out of: know which of two claims you're actually making. “We didn't have time to test it” is an honest, sometimes correct call. “We tested it and it worked” is a claim about the world, and a pre/post number without a comparison group hasn't earned the right to make it.
The size of the bet is what should set the bar, not the size of the team's confidence in it. A $500 email subject line test can run on gut feel and a coin flip's worth of consequence. A $10M feature, product line, or campaign is exactly the decision where the gap between “a number moved” and “we know why it moved” is most expensive to get wrong, and least likely to get corrected once the story everyone tells internally has already settled into “it worked.”