Every product team, marketer, and policy office eventually says some version of the same sentence: “this worked.” What that sentence rests on varies enormously, from a genuinely randomised trial to two numbers that happened to move together on a dashboard. This is a field guide to telling the difference: a ranked ladder of causal evidence, framed around product, commercial, and digital decisions, but built on the same tools economists use to study credit risk and public policy. It ends with the specific mistakes that sink careful-looking analysis, with real cases where getting this wrong (or right) changed the answer.
The story in four parts
A randomised trial and a dashboard correlation get described with the same word: evidence.
Raw correlation, matched comparison, natural experiment, randomised trial, in that order.
Right before a budget decision, when nobody's checking which rung the evidence is standing on.
Real cases below: paid search, ad measurement, credit scoring, survivorship bias.
All four of these get described, casually, as “evidence that it worked.” They are not the same evidence. Read from the top: each rung down removes one more safeguard against a confound quietly doing the work you're crediting to X.
People (or accounts, or stores) are randomly split into treatment and control, so only the intervention differs between the two groups, on average, including on things nobody thought to measure.
Feature A/B tests · marketing lift tests · a randomised pilot of a new loan-approval rule
A real-world rule, cutoff, or accident of timing splits people as if at random, without anyone designing an experiment: regression discontinuity at an eligibility threshold, difference-in-differences around a policy change, an instrument that shifts exposure for reasons unrelated to the outcome.
A price change rolled out region by region · a credit-limit rule that only applies above a score cutoff · a platform outage that hits some users and not others
Treated and untreated groups are matched or statistically adjusted to look similar on the traits you thought to record. Nobody was randomly assigned to anything, so any trait you didn't think to record can still be doing the real work.
A credit scorecard trained on approved applicants, matched to similar-looking rejected ones · a marketing report comparing “similar” cohorts before and after a campaign
Two things moved together in the data. Which one, if either, caused the other, and whether a third thing caused both, is still completely unknown.
Power users click feature X more often · most dashboard “insights” · most growth-hack case studies posted online
Moving up a rung doesn't mean the lower rungs are worthless. A correlation is a reasonable place to start looking. It's a bad place to stop, especially right before a budget decision.
None of the phrases below are lies. They're just quieter about which rung they're actually standing on than the sentence around them implies.
“Users who did X converted Y% more.”
A correlation, stated with no comparison group and no mention of who chose to do X in the first place.
“We saw a lift after launching X.”
A before-and-after on one group, with nothing else that changed over the same period ruled out.
“Our top performers all do X.”
Survivorship bias: the people who tried X and failed usually aren't in the sample being described.
“The data suggests a strong link.”
Correlational language, carrying unearned causal confidence by the time it reaches a decision-maker.
This ladder isn't a product-analytics invention. It's the ordinary hierarchy economists have used for decades to study labour markets, credit, and public policy, usually under the banner economists call the “credibility revolution”: a decades-long shift towards designs that can actually support a causal claim, not just a correlational one.
Why compare fast-food jobs on two sides of a state line, instead of just before and after?
When New Jersey raised its minimum wage in 1992, most economists expected fewer fast-food jobs to follow. Card and Krueger compared employment at fast-food restaurants in New Jersey against restaurants just across the Pennsylvania border, whose wage floor hadn't changed, using the neighbouring state as a stand-in control group for whatever else was happening in the regional economy at the time.
What it teaches: A before-and-after comparison in one place can't separate the policy's effect from everything else that also changed over that period. Finding a comparable place where the policy didn't happen turns a weak before/after into a genuine natural experiment.
How do you judge a lending rule on the applicants it never approved?
A lender only observes repayment behaviour for the applicants it actually approved. Everyone it rejected has no outcome to learn from, so a model trained naively on approved-only data is being trained on a non-random, already-filtered slice of the applicant pool. The credit-risk literature has a name for correcting this: reject inference.
What it teaches: Selection isn't only about who chooses to participate. Sometimes the selection happens on the institution's side, and the missing group is the one you'd most need to see to know if the current rule is actually working.
Does a result from a controlled lab travel to the market it's meant to describe?
This site covers the lab-versus-field gap, and the related question of whether a small pilot's effect survives being scaled up, in its own dedicated report rather than repeating it here.
What it teaches: Randomisation tells you the comparison is fair. It doesn't automatically tell you the setting, the sample, or the scale is the one you actually care about, that's a separate question, covered in full at the link below.
Every case below used real data, competent analysts, and a plausible story. Each one still produced a materially wrong answer, right up until it was checked against a design higher up the ladder.
What happens when the people who click your ad were already going to buy?
Standard marketing measurement credits a sale to whichever ad a customer clicked last. On that logic, eBay's brand-keyword paid search looked highly profitable, since people who clicked those ads converted at a high rate. The confound: people already searching for “eBay” by name were disproportionately people who were going to visit and buy anyway, ad or no ad.
What it teaches: A large-scale randomised experiment, turning brand-keyword search ads off entirely across many U.S. markets while leaving them on in others, found the true incremental effect of that spend was close to zero. The correlational read wasn't wrong about who clicked. It was wrong about what caused the sale.
How far off is the ad-effectiveness number you already trust?
Most digital ad platforms measure lift by comparing people who were shown an ad against people who weren't, using whoever was already exposed or not through the normal delivery process, not a random split. Researchers at Facebook ran large randomised experiments (holding back a genuine control group with “ghost ads”) alongside the standard observational methods, on the same campaigns, to see how far apart the two answers were.
What it teaches: The observational methods overstated the ad's true effect by a wide margin, in both directions depending on the method, because who a platform's delivery algorithm chooses to show an ad to is never a random sample of the audience. The gap wasn't a rounding error. It was large enough to change which campaigns looked worth funding.
What's missing from a case study of companies that succeeded?
This site has a full principle article on survivorship bias, the pattern behind “successful founders all did X” claims that quietly ignore every founder who did the same X and failed.
What it teaches: The fix isn't a different statistical test. It's asking who's missing from the dataset before trusting a pattern found only among the survivors.