Special Report

How to actually read behavioural science

Every principle on this site is backed by a citation, but a citation isn't a guarantee. Psychology and behavioural economics went through a real reckoning in the 2010s: some of the field's most famous, most-cited findings didn't hold up when other labs tried to reproduce them. That doesn't mean the field is broken; it means reading a study well takes more than trusting its headline finding. Two questions matter most: did this effect actually replicate under a fair, adequately powered test (internal validity), and does it hold up beyond the exact people, place, and context it was studied in (external validity)? This page walks through both, with real cases, including some famous names that don't hold up as well as they used to, and some well-intentioned interventions that made things measurably worse. For the companion question, not just whether a finding replicates, but how strong the underlying causal claim was in the first place, see How Strong Is That Evidence? A Field Guide to Causal Claims.

The replication crisis

The “replication crisis” refers to a period, especially in the 2010s, when researchers began systematically trying to reproduce classic findings, often with larger samples and pre-registered designs that couldn't be quietly adjusted after seeing the results, and found that a surprising number of famous effects shrank dramatically or vanished. It reshaped how good researchers, and good readers, evaluate a claim. Four real examples, and what each specifically teaches:

Replication

Power Posing

Can standing like a superhero for two minutes really change your hormones?

The original 2010 study, based on 42 participants, found that holding an expansive “power pose” for two minutes raised testosterone, lowered cortisol, and increased self-reported feelings of power, and it became one of the most famous findings in behavioural science, boosted by a TED talk viewed over 20 million times. A 2015 replication with a much larger, adequately powered sample (200 participants) reproduced the self-reported feeling of power, but found no effect on hormones or risk-taking behaviour at all.

What it teaches: Statistical power matters as much as the finding itself. A small original sample can produce a result that looks clean and clear but is really just noise. One of the original co-authors, Dana Carney, has since said publicly she no longer believes the hormonal and behavioural effects are real.

Carney, D. R., Cuddy, A. J. C., & Yap, A. J. (2010). “Power Posing: Brief Nonverbal Displays Affect Neuroendocrine Levels and Risk Tolerance.” Psychological Science, 21(10), 1363–1368. Replication: Ranehill, E., et al. (2015). Psychological Science, 26(5), 653–656
Replication

Ego Depletion

Why did 23 labs together find nothing, after decades of studies said willpower runs out?

For nearly two decades, one of self-control research's central ideas was that willpower works like a muscle, using it on one task temporarily depletes a limited resource, making the next act of self-control harder. Dozens of studies and a supportive meta-analysis backed it. In 2016, 23 independent labs ran the same pre-registered protocol on a combined 2,141 participants, and found, combined across all labs, no meaningful ego-depletion effect at all.

What it teaches: A large body of supporting literature isn't proof on its own. If smaller, underpowered studies pile up over years, helped by a “file-drawer problem,” where unpublished null results simply never get submitted, a real-looking scientific consensus can form around an effect that's much weaker than it appears, or isn't there. Pre-registration exists specifically to prevent this kind of drift.

Hagger, M. S., Chatzisarantis, N. L. D., et al. (2016). “A Multilab Preregistered Replication of the Ego-Depletion Effect.” Perspectives on Psychological Science, 11(4), 546–573
Replication

Facial Feedback Hypothesis

Why did the pen-in-the-teeth effect vanish the moment anyone was watching on video?

The original “pen in the mouth” study found people rated cartoons funnier when a pen held in their teeth forced their face into a smile, versus a pen held in their lips forcing a pout, a clean demonstration that an expression can shape the emotion it's usually thought to just express. A 2016 replication across 17 labs and over 2,000 participants, following a protocol vetted by the original author, found no effect at all.

What it teaches: External validity can hinge on a procedural detail nobody thought to flag. One proposed explanation: the replication labs, unlike the original, recorded participants on video, and being watched appears to suppress the effect being measured. A failed replication doesn't always mean the original effect was fake. Sometimes it means the effect has a boundary condition nobody had identified yet.

Strack, F., Martin, L. L., & Stepper, S. (1988). Journal of Personality and Social Psychology, 54(5), 768–777. Replication: Wagenmakers, E.-J., et al. (2016). Perspectives on Psychological Science, 11(6), 917–928
Replication

Behavioural Priming

Why did people only “walk slower” after old-age words when the experimenter expected it?

The original, hugely influential study found participants exposed to words associated with old age (“Florida,” “wrinkle,” “bingo”) walked measurably slower leaving the lab afterwards. A 2012 replication using objective infrared timing, rather than a human experimenter with a stopwatch, found no priming effect in its first experiment. In a second experiment, the effect only appeared when the experimenters running the stopwatch themselves believed, in advance, that participants would walk slower.

What it teaches: A textbook internal-validity failure, not a fake effect exactly, but a real effect coming from the wrong cause. If whoever measures the outcome knows or can guess the hypothesis and isn't blinded to it, their own expectations can leak into a subjective judgment like “how slowly did that person seem to walk.” Objective, blinded measurement is what internal validity is protecting.

Bargh, J. A., Chen, M., & Burrows, L. (1996). Journal of Personality and Social Psychology, 71(2), 230–244. Replication: Doyen, S., et al. (2012). “Behavioral Priming: It's All in the Mind, but Whose Mind?” PLOS ONE, 7(1), e29081
Internal, and four more: Five Ways an Experiment Can Be Right and Still Wrong covers this exact failure mode alongside external, ecological, construct, and statistical conclusion validity, the four questions this page's own “internal vs. external” framing doesn't ask on its own.

The throughline across all four: a single study, however clean it looks, is a data point, not a verdict. Before treating any single finding, including ones on this site, as settled, it's worth asking: was the sample large enough to detect a real effect reliably? Was the analysis locked in before the results came in, or could it have been adjusted after the fact? Has anyone else, in a different lab with different participants, got the same result? And does the population and setting it was tested in actually resemble the one you're trying to apply it to? None of that means dismissing behavioural science. It means reading it with the methodology in view, not just the headline.

When interventions backfire

A finding replicating in a lab is still one step removed from an intervention working as intended once it's deployed. Even a genuinely real, well-replicated mechanism can produce the opposite of its intended effect once it's applied inside a messy, real-world system with its own incentives and substitutes. Six real cases, six different mechanisms, now have their own full report.

Special Report

Backfires have six different causes. They all look the same.

A fine, a ban, a scare tactic, a comparison letter, a bounty, and a reminder email each produced the exact opposite of what it was built to do. Crowding out, substitution, a wrong theory of the audience, a hidden subgroup, a gamed proxy, and an uncounted second number, six specific, nameable mechanisms, not one shapeless category called “it backfired.”

Lab vs. field: why where you test it matters

A finding demonstrated in a controlled lab and the same finding holding up in a messy real market are not the same evidence, even when they're described with the same word: “effect.” Economist John List has spent much of his career specifically probing that gap, and the honest answer is that it's often larger than people assume.

Validity

Lab Generosity Doesn't Always Travel

Why doesn't how you play a lab game predict how generous you actually are?

Many well-known findings about generosity and fairness, such as how much people give in a dictator game or how they punish unfair offers, come from lab experiments with paid student volunteers who know they're being watched. List and co-author Steven Levitt reviewed the evidence and found these lab measures frequently fail to predict how the same people behave in comparable real-world situations, where anonymity, social context, and real stakes differ from the lab.

What it teaches: A controlled setting buys precision, but it can quietly change the very behaviour being measured: being watched, knowing it's an experiment, and interacting with strangers you'll never see again are all real features of a lab that don't exist in the market the finding is meant to describe.

Levitt, S. D., & List, J. A. (2007). “What Do Laboratory Experiments Measuring Social Preferences Reveal About the Real World?” Journal of Economic Perspectives, 21(2), 153–174
Validity

The Voltage Drop at Scale

Why do pilots that work beautifully shrink the moment you roll them out for real?

Across List's own applied research programme, a recurring pattern shows up: an intervention that produces a strong effect in a small pilot frequently produces a much smaller one, sometimes a fraction the size, once rolled out at real scale. He calls this the “voltage drop.” Worth flagging honestly: this specific framing comes from List's popular synthesis book rather than one single peer-reviewed paper, so treat it as an experienced practitioner's pattern-observation, not a controlled finding in its own right.

What it teaches: A successful small pilot is genuinely useful evidence, but it isn't a scale-tested one. The population, the staff running it, and the novelty of a small trial can all inflate results in ways that don't survive contact with a full rollout, worth checking for, not assuming away.

List, J. A. (2022). The Voltage Effect: How to Make Good Ideas Great and Great Ideas Scale. Currency, drawing on List's broader field-experiment research programme, cited throughout in peer-reviewed form.
The formal name for this: External validity: does it hold up anywhere but here?, on Five Ways an Experiment Can Be Right and Still Wrong, including the 17-lab replication that pinned down exactly why a famous facial-feedback effect only worked when nobody was watching.

Why asking people what they'll do is unreliable

Surveys asking “would you pay for this?” or “would you do this?” are a normal, useful research tool, but they measure something different from real behaviour, and the gap between the two has its own name.

Validity

Cheap Talk & Hypothetical Bias

Why do people say they'd pay more than they actually do, even when warned about it?

When there's no real cost to answering however makes you look good, people reliably state a higher willingness to pay in hypothetical surveys than they actually pay when real money is on the line, a well-documented gap known as hypothetical bias. Attempts to fix it by simply warning respondents about the bias beforehand (“cheap talk” scripts) have generally proven weak or unreliable at closing the gap.

What it teaches: A stated intention is not a prediction of behaviour, even from an honest, well-meaning respondent. It's an answer to a different, lower-stakes question. Wherever possible, real behaviour (an actual purchase, an actual signed commitment) is a stronger form of evidence than a stated one, and this site tries to flag the difference rather than blur it.

Harrison, G. W., & Rutström, E. E. (2008). “Experimental Evidence on the Existence of Hypothetical Bias in Value Elicitation Methods,” in Handbook of Experimental Economics Results, Vol. 1. Elsevier.
The formal name for this: Ecological validity: does the task even resemble the real thing?, on Five Ways an Experiment Can Be Right and Still Wrong, including this exact study and this site's own Risk Aversion entry on real versus hypothetical stakes.

More special reports

Nineteen companion pieces on evidence quality, research honesty, and applied mechanism design, written in the same spirit as this page.

Field Guide

How strong is that evidence, actually?

A ranked ladder of causal evidence, from a genuine randomised trial down to two numbers that happened to move together on a dashboard, so a claim of “this worked” can actually be checked.

Special Report

Policy-based evidence, not evidence-based policy

Economist John List’s case that most cited evidence comes from a small pilot never designed to survive being scaled up, and what testing for scale before rollout looks like instead.

Special Report

Don’t say the E-word

Why so many organisations will run a test but flinch at calling it an experiment, and what that flinch quietly costs a team’s ability to learn from what it ships.

Special Report

How you pay changes what you spend

Two real studies, cash versus debit and cash versus credit, and the shared mechanism, payment transparency, that connects them into one honest gradient.

Special Report

The test was real. The conclusion wasn’t.

Thirteen real experiments that misled without lying: samples too small to see the effect, results stretched from one place into a universal law, and tests run on a guess with no theory behind the metric they moved.

Special Report

The lottery you can’t lose

Prize-linked savings accounts trade a guaranteed rate for a lottery-style prize, without ever risking the deposit itself. The mechanism, the real programmes, and the honest catch.

Special Report

Behavioural economics vs. behavioural science vs. psychology vs. UX

Behavioural economics, psychology, applied behavioural science, and UX all draw on the same small set of findings. Two real, traceable examples of what happens when a name survives the trip between them, and what happens when it doesn’t.

Special Report

Six moves behind every pricing page

Temporal reframing, chunking, anchoring, bundling, sunk cost, and zero price: six real, cited mechanisms, and the pricing-page patterns that run on each one, none of them requiring deception.

Special Report

Not testing is still a bet

A pre-vs-post number isn't a smaller version of measurement, it usually isn't measurement at all. Real cases where it said one thing and a real experiment said another, eBay's paid search among them.

Special Report

Every default decides who pays for doing nothing

Opt-in, opt-out, forced choice, timed defaults, and personalised defaults are five separate levers. The real studies behind each setting, and the famous organ-donation result that didn’t survive a better test.

Special Report

The fine print nobody reads

Almost nobody opens a standard-form disclosure before agreeing to it, and when a conflict-of-interest disclosure is read, it can make the advice behind it worse, not better. The one disclosure format a real field test found actually changed behaviour.

Special Report

You're already doing behavioural economics

Anchoring, the decoy effect, defaults, scarcity, social proof, and four more: eight mechanisms already running on any pricing page or product screen, whether or not whoever built it could name them.

Special Report

Proximity gets mistaken for merit

A 1946 MIT housing lottery and a 2022 machine-learning study of the MLB draft, seventy years apart, found the identical bias: living close to a decision-maker predicts the outcome, even when it has nothing to do with merit.

Special Report

Printing costs more than digital. The paper copy still wins.

Digital is cheaper to send, every time. Three real, cited studies on touch, reading comprehension, and how people value physical objects explain why organisations, NRMA's own quarterly member magazine among them, still mail the paper copy anyway.

Special Report

Breaking a form into steps changes how it feels. It doesn’t always change how it performs.

Progress Flow (steps and a progress bar) and Progressive Load (one continuous page) both borrow real psychology, chunking and the goal-gradient effect, but six studies ranked by fidelity show the same pattern doing nothing on a desktop and backfiring on a phone.

Special Report

AI gets over-trusted until it errs once. Then it gets under-trusted for good.

Automation bias and algorithm aversion are the same miscalibration pointed in opposite directions, and keeping a human in the loop doesn’t reliably fix either one. Ten real, cited studies across AI coding assistants, clinical AI, and everyday knowledge work.

Special Report

The coin has no memory. The streak still feels real.

What a fair coin can teach you about statistics: why a real random sequence looks streakier than expected, why a run of heads doesn't make tails overdue, what a p-value actually says, and why the smallest sample in any ranking always looks the most extreme.

Special Report

The biases draining your super could also fill it

Ten million unintended duplicate super accounts, tens of thousands of members who sold at the exact bottom of the 2020 crash, and $37.8 billion pulled out early during the pandemic. Four real, cited mechanisms, and the one 2021 law change that turned the same default that caused the damage into the fix.

Special Report

Five ways an experiment can be right and still wrong

Internal, external, ecological, construct, and statistical conclusion validity: five separate questions a single result can pass or fail independently, each demonstrated on a real, cited replication case already documented on this page.

Special Report

Your roadmap has a Phase 2. Most skip straight past it.

The outcome chain behind a feature launch, mapped stage by stage on one worked example: awareness, discovery, trial, onboarding, completion, and whether it actually persists, with the real proxy trap waiting at every link.

Special Report

Every program runs on a theory. Most never write it down.

The decades-old evaluation discipline underneath every logic model: why a program fails for one of two very different reasons, shown on a real program that made crime worse while working exactly as designed, and one that got the theory right from day one.