Special Report

Don't say the E-word

Economist John List has run field experiments with hundreds of organisations, and one small piece of advice he gives researchers trying to get started keeps resurfacing: never call it an experiment. Call it a pilot, or a trial. Many organisations hear the word “experiment” and react as though it's slightly improper, even when they're entirely willing to run the identical test under a friendlier name. The word itself isn't really the problem. It's a visible symptom of something more costly: an organisation avoiding the actual discipline, a real control group, a fair comparison, a willingness to be shown wrong, that the word implies.

The story in four parts

Organisations flinch at the word “experiment”

List's own advice to researchers: call it a pilot or a trial instead, many partners hear “experiment” as improper.

The word goes, and the control group often goes with it

A “pilot” that launches to everyone at once, with no comparison group, isn't a gentler experiment. It's a lower rung of evidence.

The real cost is what you don't find out

Most teams weigh the cost of testing. Few weigh the cost of never knowing whether the untested version actually works.

The organisations willing to say it plainly find out more

Uber's own apology experiment, and List's early-childhood trial in Chicago Heights, both happened because someone was willing to call it what it was.

The evidence at a glance

A real field economist's own Rule #1

List's practitioner guide for researchers approaching organisations: drop the word “experiment” entirely and call it a pilot or a trial instead.

A repeated apology can backfire

A real field experiment on customer apologies found repeating the identical apology on a second failure backfired, the opposite of the obvious fix.

Governments avoid the word most

List and Heffetz describe governments as unusually resistant to running randomised trials, even relative to companies that A/B test routinely.

The word hides at least six different methods

Part of why “experiment” spooks people is that it doesn't mean one thing. A researcher, a product manager, and a UX designer can all use the word in the same meeting and picture three different methods, with three very different levels of rigor behind them. Sorting out which one is actually on the table matters more than the label itself.

Economists Glenn Harrison and John List, the same List behind this report's own citation, published an influential framework for sorting real experiments by exactly this: how real the subjects, the stakes, and the setting actually are. It runs from a conventional lab experiment (a researcher-built setting, a fixed pool of paid volunteers) to a natural field experiment (real people doing what they'd do anyway, with no idea they're in a study at all). Everything else in the field, a five-user usability test, a live A/B test, a hypothetical survey question, sits somewhere on that same spectrum, whether or not the person running it thinks of it that way.

Harrison, G. W., & List, J. A. (2004). “Field Experiments.” Journal of Economic Literature, 42(4), 1009–1055.
Control over thecomparison ↑ Real-world stakes & realism → Hypothetical scenarios stated intent, not real behaviour Lab experiments controlled, but artificial User testing diagnostic, no control group RCTs, A/B tests included real stakes, real randomisation Multivariate tests more factors, thinner power Natural experiments real world, control not guaranteed

Realism and control usually trade off against each other, and randomised field experiments, the RCTs and A/B tests in the upper right, are the rare design that gets both at once. Everything else on the chart is paying for realism with less certainty, or paying for certainty with less realism.

Lowest realism

Hypothetical / imagined-scenario studies

Ask people what they'd do, or pay, in a scenario they only imagine, never actually live through.

Strength: cheap, fast, and the only option for a situation too risky or too far off to actually build, like a product that doesn't exist yet.

Weakness: what people say they'd do and what they actually do reliably diverge. A meta-analysis of 29 studies found hypothetical willingness-to-pay and willingness-to-accept overstated real, incentivised values by roughly a factor of three, on average (List & Gallet, 2001).

Highest control

Lab experiments

Randomly assign real participants to conditions inside a setting the researcher builds, isolating one variable at a time.

Strength: the tightest possible control over confounds. If the lab result shows an effect, it's very unlikely to be a coincidence.

Weakness: the setting, the incentives, and often the subject pool are artificial. A result that holds in a lab doesn't guarantee it survives real stakes and real distractions.

Diagnostic, not comparative

User testing / usability testing

A handful of real users attempt real tasks on the real product while someone watches, usually thinking aloud.

Strength: fast and rich. Jakob Nielsen's own model puts five users at catching roughly 85% of usability problems (Nielsen, 2000), at a fraction of the cost of a bigger study.

Weakness: no control group and no counterfactual. It shows where people got stuck, not whether version A actually converts more people than version B, or by how much.

Realism and control together

Randomised Controlled Trials, A/B tests included

Randomly split real people into two or more real conditions and measure a real outcome. “A/B test” is simply what a digital product team calls the identical design run inside a live app or website.

Strength: random assignment plus real stakes is the combination that earns the name “gold standard”, high confidence a measured difference is actually caused by the change.

Weakness: needs enough real traffic or participants to detect the effect at all (see this site's own sample-size guidance). Testing one change at a time is also slow when there are many candidates to try.

Same design, more factors

Multivariate tests

The same randomised design as an A/B test, but applied to several factors at once: a headline, an image, a price display, tested together instead of one change at a time.

Strength: reveals interaction effects a sequence of single-factor A/B tests would miss entirely, like a headline that only works with one particular image.

Weakness: every added factor multiplies the number of combinations, so the same total traffic gets split thinner across more cells, and each individual comparison earns less statistical certainty.

Highest realism

Natural experiments / quasi-experiments

No one designs the comparison. A real policy change, price change, or eligibility cutoff creates something close to random variation on its own, and the researcher finds it after the fact.

Strength: the highest possible realism, real people making real decisions, with zero risk of the study itself changing anyone's behaviour.

Weakness: whether the comparison is really “as good as random” is an assumption, not a guarantee, and the whole result rests on how well that assumption holds. See Natural Experiments for the full page on how that assumption gets tested.

Why the word itself is the barrier

List sets this out plainly in his own practitioner guide to running field experiments: many organizational partners treat the word “experiment” as close to repugnant, so the practical advice to a researcher trying to start a collaboration is to drop the word entirely and use “pilot” or “trial” instead. He has repeated the same point publicly and more casually: in a 2022 post, he called it “Rule #1” for researchers approaching organisations, and linked it to a companion piece he co-wrote on why governments in particular resist evidence-based policymaking.

Read together, the peer-reviewed version and the public restatement are making the same point at two different distances. The academic paper documents the pattern as professional advice, tested across List's own long career of setting up organizational field experiments. The public post is List saying the same thing plainly, to a wider audience, eleven years later. Treat the paper as the load-bearing citation and the post as the accessible restatement of it, not as two independent pieces of evidence.

List, J. A. (2011). “Why Economists Should Conduct Field Experiments and 14 Tips for Pulling One Off.” Journal of Economic Perspectives, 25(3), 3–16. Restated publicly: List, J. A. (2022, February). Rule #1: never use the “E” word. Referencing Heffetz, O., & List, J. A. (2021). “Who's Afraid of Evidence-Based Policymaking?

Why avoiding the word costs more than a word

The word swap alone would be harmless if the underlying test stayed the same. The actual risk is that avoiding “experiment” often travels together with avoiding the thing that makes an experiment worth running in the first place: a genuine control group, and a willingness to measure an outcome that might come back disappointing. A “pilot” that launches to everyone at once, with no comparison group and no pre-registered way to fail, isn't a gentler version of an experiment. On the evidence ladder this site sets out separately, it's sitting on a much lower rung, no matter what it's called.

List's other practical point about organizational resistance reframes where the real cost sits. The conversation inside most organisations focuses on the cost of running a test (the time, the resourcing, the risk of a visible failure), and rarely on the opportunity cost of not running one. That opportunity cost is continuing to fund or scale a programme nobody has actually measured against a fair alternative. Once that opportunity cost is made explicit, the case for testing gets much harder to talk yourself out of.

The same avoidance, in commercial and product terms

This site's own “Not Testing Is Still a Bet” report covers the product and commercial version of the identical avoidance: teams that ship an untested change to everyone at once because running a proper test on real users feels like it's the riskier, less ethical choice. The reverse is usually true, and the pre/post number they'll use to call it a win afterwards isn't measuring anything either.

Cross-link

Not Testing Is Still a Bet

Which is actually riskier: testing a change on some users, or shipping it to everyone untested?

This site's full report on the subject argues the second option is usually treated as the “safe” default, and usually isn't, then shows real cases where the number a team trusted to say it worked was never actually measuring that.

What it teaches: “We didn't want to experiment on our users” is almost always a decision to skip measurement, not a decision to avoid risk. See the full report for the reasoning and the real cases.

See Not Testing Is Still a Bet for the full report.
Cross-link

What Testing Actually Bought, Once Run

What did a company find once it stopped guessing and ran the test anyway?

This site's experiment blueprint on costly, specific apologies builds directly on List's own field experiment with Uber: a real organisation willing to run the test, and to be shown a genuinely counterintuitive result (a repeated identical apology can backfire) rather than only the flattering one.

What it teaches: The organisations willing to say “experiment” and mean it are the ones who end up finding out things a pilot dressed up as a done deal never would.