Field Guide

How strong is that evidence, actually?

Every product team, marketer, and policy office eventually says some version of the same sentence: “this worked.” What that sentence rests on varies enormously, from a genuinely randomised trial to two numbers that happened to move together on a dashboard. This is a field guide to telling the difference: a ranked ladder of causal evidence, framed around product, commercial, and digital decisions, but built on the same tools economists use to study credit risk and public policy. It ends with the specific mistakes that sink careful-looking analysis, with real cases where getting this wrong (or right) changed the answer.

The story in four parts

Every “it worked” claim rests on different ground

A randomised trial and a dashboard correlation get described with the same word: evidence.

Rank them, weakest to strongest

Raw correlation, matched comparison, natural experiment, randomised trial, in that order.

The weakest kind gets funded most often

Right before a budget decision, when nobody's checking which rung the evidence is standing on.

Here's the full ladder, and where it breaks

Real cases below: paid search, ad measurement, credit scoring, survivorship bias.

The evidence ladder: four ways to test whether X caused Y

All four of these get described, casually, as “evidence that it worked.” They are not the same evidence. Read from the top: each rung down removes one more safeguard against a confound quietly doing the work you're crediting to X.

Randomised controlled trial
Strongest

People (or accounts, or stores) are randomly split into treatment and control, so only the intervention differs between the two groups, on average, including on things nobody thought to measure.

Feature A/B tests · marketing lift tests · a randomised pilot of a new loan-approval rule

Individual RCTEach person or account is randomised on their own, the classic feature A/B test.
Cluster RCTWhole stores, branches, or regions are randomised together, used when a treatment can't be applied to one person at a time.
Stepped-wedge rolloutEveryone eventually gets the treatment; only the order of rollout is randomised.
Natural experiment
High

A real-world rule, cutoff, or accident of timing splits people as if at random, without anyone designing an experiment: regression discontinuity at an eligibility threshold, difference-in-differences around a policy change, an instrument that shifts exposure for reasons unrelated to the outcome.

A price change rolled out region by region · a credit-limit rule that only applies above a score cutoff · a platform outage that hits some users and not others

Regression discontinuityCompares people just above and just below a real cutoff, a credit score threshold, an age limit, who are otherwise near-identical.
Difference-in-differencesCompares the change over time in a group affected by a policy against the change in a similar, unaffected group.
Instrumental variableUses a third factor that shifts exposure to the treatment for reasons unrelated to the outcome, to isolate its effect.
Matched observational comparison
Moderate

Treated and untreated groups are matched or statistically adjusted to look similar on the traits you thought to record. Nobody was randomly assigned to anything, so any trait you didn't think to record can still be doing the real work.

A credit scorecard trained on approved applicants, matched to similar-looking rejected ones · a marketing report comparing “similar” cohorts before and after a campaign

Propensity score matchingStatistically pairs treated and untreated units that look alike on measured characteristics.
Regression with controlsAdjusts for known confounders inside a single statistical model, rather than physically matching pairs.
Raw correlation
Lowest

Two things moved together in the data. Which one, if either, caused the other, and whether a third thing caused both, is still completely unknown.

Power users click feature X more often · most dashboard “insights” · most growth-hack case studies posted online

Bivariate correlationTwo variables measured and compared, with no other factors accounted for.
Anecdote or single caseOne example treated as if it were representative of a general pattern.

Moving up a rung doesn't mean the lower rungs are worthless. A correlation is a reasonable place to start looking. It's a bad place to stop, especially right before a budget decision.

Hallmarks: what a lower rung sounds like when it's dressed up as a higher one

None of the phrases below are lies. They're just quieter about which rung they're actually standing on than the sentence around them implies.

“Users who did X converted Y% more.”

A correlation, stated with no comparison group and no mention of who chose to do X in the first place.

“We saw a lift after launching X.”

A before-and-after on one group, with nothing else that changed over the same period ruled out.

“Our top performers all do X.”

Survivorship bias: the people who tried X and failed usually aren't in the sample being described.

“The data suggests a strong link.”

Correlational language, carrying unearned causal confidence by the time it reaches a decision-maker.

Four questions that do most of the work

  1. Compared to what?Is there an actual comparison group, or just a before-and-after on the same people?
  2. Who's missing, and why?Could selection into or out of the sample, who clicked, who was approved, who's still in business, be driving the result?
  3. Is this the headline result, or a secondary point being promoted?A true statement pulled from deep in a report can still misrepresent what the study actually found first.
  4. Would this survive real scale, with ordinary staff and the full population?A result from a small, favourably-run pilot answers a narrower question than the one usually being asked of it.

The same ladder, outside product and marketing

This ladder isn't a product-analytics invention. It's the ordinary hierarchy economists have used for decades to study labour markets, credit, and public policy, usually under the banner economists call the “credibility revolution”: a decades-long shift towards designs that can actually support a causal claim, not just a correlational one.

Natural experiment

The Minimum-Wage Border Study

Why compare fast-food jobs on two sides of a state line, instead of just before and after?

When New Jersey raised its minimum wage in 1992, most economists expected fewer fast-food jobs to follow. Card and Krueger compared employment at fast-food restaurants in New Jersey against restaurants just across the Pennsylvania border, whose wage floor hadn't changed, using the neighbouring state as a stand-in control group for whatever else was happening in the regional economy at the time.

What it teaches: A before-and-after comparison in one place can't separate the policy's effect from everything else that also changed over that period. Finding a comparable place where the policy didn't happen turns a weak before/after into a genuine natural experiment.

Card, D., & Krueger, A. B. (1994). “Minimum Wages and Employment: A Case Study of the Fast-Food Industry in New Jersey and Pennsylvania.” American Economic Review, 84(4), 772–793
Field study type

Credit Scoring's Missing Half

How do you judge a lending rule on the applicants it never approved?

A lender only observes repayment behaviour for the applicants it actually approved. Everyone it rejected has no outcome to learn from, so a model trained naively on approved-only data is being trained on a non-random, already-filtered slice of the applicant pool. The credit-risk literature has a name for correcting this: reject inference.

What it teaches: Selection isn't only about who chooses to participate. Sometimes the selection happens on the institution's side, and the missing group is the one you'd most need to see to know if the current rule is actually working.

Hand, D. J., & Henley, W. E. (1997). “Statistical Classification Methods in Consumer Credit Scoring: A Review.” Journal of the Royal Statistical Society: Series A, 160(3), 523–541
Field study type

Lab vs. Field, Again

Does a result from a controlled lab travel to the market it's meant to describe?

This site covers the lab-versus-field gap, and the related question of whether a small pilot's effect survives being scaled up, in its own dedicated report rather than repeating it here.

What it teaches: Randomisation tells you the comparison is fair. It doesn't automatically tell you the setting, the sample, or the scale is the one you actually care about, that's a separate question, covered in full at the link below.

See Reading the Research: Lab vs. field, Why Governments Need Policy-Based Evidence for the scaling version of this problem, and Five Ways an Experiment Can Be Right and Still Wrong for this question's formal name, external validity, alongside four others.

Where careful-looking analysis still goes wrong

Every case below used real data, competent analysts, and a plausible story. Each one still produced a materially wrong answer, right up until it was checked against a design higher up the ladder.

Selection bias

The Paid Search That Wasn't Doing Much

What happens when the people who click your ad were already going to buy?

Standard marketing measurement credits a sale to whichever ad a customer clicked last. On that logic, eBay's brand-keyword paid search looked highly profitable, since people who clicked those ads converted at a high rate. The confound: people already searching for “eBay” by name were disproportionately people who were going to visit and buy anyway, ad or no ad.

What it teaches: A large-scale randomised experiment, turning brand-keyword search ads off entirely across many U.S. markets while leaving them on in others, found the true incremental effect of that spend was close to zero. The correlational read wasn't wrong about who clicked. It was wrong about what caused the sale.

Blake, T., Nosko, C., & Tadelis, S. (2015). “Consumer Heterogeneity and Paid Search Effectiveness: A Large-Scale Field Experiment.” Econometrica, 83(1), 155–174
Measurement bias

Ad Measurement, Checked Against Itself

How far off is the ad-effectiveness number you already trust?

Most digital ad platforms measure lift by comparing people who were shown an ad against people who weren't, using whoever was already exposed or not through the normal delivery process, not a random split. Researchers at Facebook ran large randomised experiments (holding back a genuine control group with “ghost ads”) alongside the standard observational methods, on the same campaigns, to see how far apart the two answers were.

What it teaches: The observational methods overstated the ad's true effect by a wide margin, in both directions depending on the method, because who a platform's delivery algorithm chooses to show an ad to is never a random sample of the audience. The gap wasn't a rounding error. It was large enough to change which campaigns looked worth funding.

Gordon, B. R., Zettelmeyer, F., Bhargava, N., & Chapsky, D. (2019). “A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook.” Marketing Science, 38(2), 193–225
Survivorship bias

Counting Only the Cases Still in the Room

What's missing from a case study of companies that succeeded?

This site has a full principle article on survivorship bias, the pattern behind “successful founders all did X” claims that quietly ignore every founder who did the same X and failed.

What it teaches: The fix isn't a different statistical test. It's asking who's missing from the dataset before trusting a pattern found only among the survivors.

See Survivorship Bias for the full mechanism and citation.