Every principle on this site is backed by a citation, but a citation isn't a guarantee. Psychology and behavioural economics went through a real reckoning in the 2010s: some of the field's most famous, most-cited findings didn't hold up when other labs tried to reproduce them. That doesn't mean the field is broken; it means reading a study well takes more than trusting its headline finding. Two questions matter most: did this effect actually replicate under a fair, adequately powered test (internal validity), and does it hold up beyond the exact people, place, and context it was studied in (external validity)? This page walks through both, with real cases, including some famous names that don't hold up as well as they used to, and some well-intentioned interventions that made things measurably worse. For the companion question, not just whether a finding replicates, but how strong the underlying causal claim was in the first place, see How Strong Is That Evidence? A Field Guide to Causal Claims.
The “replication crisis” refers to a period, especially in the 2010s, when researchers began systematically trying to reproduce classic findings, often with larger samples and pre-registered designs that couldn't be quietly adjusted after seeing the results, and found that a surprising number of famous effects shrank dramatically or vanished. It reshaped how good researchers, and good readers, evaluate a claim. Four real examples, and what each specifically teaches:
Can standing like a superhero for two minutes really change your hormones?
The original 2010 study, based on 42 participants, found that holding an expansive “power pose” for two minutes raised testosterone, lowered cortisol, and increased self-reported feelings of power, and it became one of the most famous findings in behavioural science, boosted by a TED talk viewed over 20 million times. A 2015 replication with a much larger, adequately powered sample (200 participants) reproduced the self-reported feeling of power, but found no effect on hormones or risk-taking behaviour at all.
What it teaches: Statistical power matters as much as the finding itself. A small original sample can produce a result that looks clean and clear but is really just noise. One of the original co-authors, Dana Carney, has since said publicly she no longer believes the hormonal and behavioural effects are real.
Why did 23 labs together find nothing, after decades of studies said willpower runs out?
For nearly two decades, one of self-control research's central ideas was that willpower works like a muscle, using it on one task temporarily depletes a limited resource, making the next act of self-control harder. Dozens of studies and a supportive meta-analysis backed it. In 2016, 23 independent labs ran the same pre-registered protocol on a combined 2,141 participants, and found, combined across all labs, no meaningful ego-depletion effect at all.
What it teaches: A large body of supporting literature isn't proof on its own. If smaller, underpowered studies pile up over years, helped by a “file-drawer problem,” where unpublished null results simply never get submitted, a real-looking scientific consensus can form around an effect that's much weaker than it appears, or isn't there. Pre-registration exists specifically to prevent this kind of drift.
Why did the pen-in-the-teeth effect vanish the moment anyone was watching on video?
The original “pen in the mouth” study found people rated cartoons funnier when a pen held in their teeth forced their face into a smile, versus a pen held in their lips forcing a pout, a clean demonstration that an expression can shape the emotion it's usually thought to just express. A 2016 replication across 17 labs and over 2,000 participants, following a protocol vetted by the original author, found no effect at all.
What it teaches: External validity can hinge on a procedural detail nobody thought to flag. One proposed explanation: the replication labs, unlike the original, recorded participants on video, and being watched appears to suppress the effect being measured. A failed replication doesn't always mean the original effect was fake. Sometimes it means the effect has a boundary condition nobody had identified yet.
Why did people only “walk slower” after old-age words when the experimenter expected it?
The original, hugely influential study found participants exposed to words associated with old age (“Florida,” “wrinkle,” “bingo”) walked measurably slower leaving the lab afterwards. A 2012 replication using objective infrared timing, rather than a human experimenter with a stopwatch, found no priming effect in its first experiment. In a second experiment, the effect only appeared when the experimenters running the stopwatch themselves believed, in advance, that participants would walk slower.
What it teaches: A textbook internal-validity failure, not a fake effect exactly, but a real effect coming from the wrong cause. If whoever measures the outcome knows or can guess the hypothesis and isn't blinded to it, their own expectations can leak into a subjective judgment like “how slowly did that person seem to walk.” Objective, blinded measurement is what internal validity is protecting.
The throughline across all four: a single study, however clean it looks, is a data point, not a verdict. Before treating any single finding, including ones on this site, as settled, it's worth asking: was the sample large enough to detect a real effect reliably? Was the analysis locked in before the results came in, or could it have been adjusted after the fact? Has anyone else, in a different lab with different participants, got the same result? And does the population and setting it was tested in actually resemble the one you're trying to apply it to? None of that means dismissing behavioural science. It means reading it with the methodology in view, not just the headline.
A finding replicating in a lab is still one step removed from an intervention working as intended once it's deployed. Even a genuinely real, well-replicated mechanism can produce the opposite of its intended effect once it's applied inside a messy, real-world system with its own incentives and substitutes. Six real cases, six different mechanisms, now have their own full report.
A fine, a ban, a scare tactic, a comparison letter, a bounty, and a reminder email each produced the exact opposite of what it was built to do. Crowding out, substitution, a wrong theory of the audience, a hidden subgroup, a gamed proxy, and an uncounted second number, six specific, nameable mechanisms, not one shapeless category called “it backfired.”
A finding demonstrated in a controlled lab and the same finding holding up in a messy real market are not the same evidence, even when they're described with the same word: “effect.” Economist John List has spent much of his career specifically probing that gap, and the honest answer is that it's often larger than people assume.
Why doesn't how you play a lab game predict how generous you actually are?
Many well-known findings about generosity and fairness, such as how much people give in a dictator game or how they punish unfair offers, come from lab experiments with paid student volunteers who know they're being watched. List and co-author Steven Levitt reviewed the evidence and found these lab measures frequently fail to predict how the same people behave in comparable real-world situations, where anonymity, social context, and real stakes differ from the lab.
What it teaches: A controlled setting buys precision, but it can quietly change the very behaviour being measured: being watched, knowing it's an experiment, and interacting with strangers you'll never see again are all real features of a lab that don't exist in the market the finding is meant to describe.
Why do pilots that work beautifully shrink the moment you roll them out for real?
Across List's own applied research programme, a recurring pattern shows up: an intervention that produces a strong effect in a small pilot frequently produces a much smaller one, sometimes a fraction the size, once rolled out at real scale. He calls this the “voltage drop.” Worth flagging honestly: this specific framing comes from List's popular synthesis book rather than one single peer-reviewed paper, so treat it as an experienced practitioner's pattern-observation, not a controlled finding in its own right.
What it teaches: A successful small pilot is genuinely useful evidence, but it isn't a scale-tested one. The population, the staff running it, and the novelty of a small trial can all inflate results in ways that don't survive contact with a full rollout, worth checking for, not assuming away.
Surveys asking “would you pay for this?” or “would you do this?” are a normal, useful research tool, but they measure something different from real behaviour, and the gap between the two has its own name.
Why do people say they'd pay more than they actually do, even when warned about it?
When there's no real cost to answering however makes you look good, people reliably state a higher willingness to pay in hypothetical surveys than they actually pay when real money is on the line, a well-documented gap known as hypothetical bias. Attempts to fix it by simply warning respondents about the bias beforehand (“cheap talk” scripts) have generally proven weak or unreliable at closing the gap.
What it teaches: A stated intention is not a prediction of behaviour, even from an honest, well-meaning respondent. It's an answer to a different, lower-stakes question. Wherever possible, real behaviour (an actual purchase, an actual signed commitment) is a stronger form of evidence than a stated one, and this site tries to flag the difference rather than blur it.
Nineteen companion pieces on evidence quality, research honesty, and applied mechanism design, written in the same spirit as this page.
A ranked ladder of causal evidence, from a genuine randomised trial down to two numbers that happened to move together on a dashboard, so a claim of “this worked” can actually be checked.
Economist John List’s case that most cited evidence comes from a small pilot never designed to survive being scaled up, and what testing for scale before rollout looks like instead.
Why so many organisations will run a test but flinch at calling it an experiment, and what that flinch quietly costs a team’s ability to learn from what it ships.
Two real studies, cash versus debit and cash versus credit, and the shared mechanism, payment transparency, that connects them into one honest gradient.
Thirteen real experiments that misled without lying: samples too small to see the effect, results stretched from one place into a universal law, and tests run on a guess with no theory behind the metric they moved.
Prize-linked savings accounts trade a guaranteed rate for a lottery-style prize, without ever risking the deposit itself. The mechanism, the real programmes, and the honest catch.
Behavioural economics, psychology, applied behavioural science, and UX all draw on the same small set of findings. Two real, traceable examples of what happens when a name survives the trip between them, and what happens when it doesn’t.
Temporal reframing, chunking, anchoring, bundling, sunk cost, and zero price: six real, cited mechanisms, and the pricing-page patterns that run on each one, none of them requiring deception.
A pre-vs-post number isn't a smaller version of measurement, it usually isn't measurement at all. Real cases where it said one thing and a real experiment said another, eBay's paid search among them.
Opt-in, opt-out, forced choice, timed defaults, and personalised defaults are five separate levers. The real studies behind each setting, and the famous organ-donation result that didn’t survive a better test.
Almost nobody opens a standard-form disclosure before agreeing to it, and when a conflict-of-interest disclosure is read, it can make the advice behind it worse, not better. The one disclosure format a real field test found actually changed behaviour.
Anchoring, the decoy effect, defaults, scarcity, social proof, and four more: eight mechanisms already running on any pricing page or product screen, whether or not whoever built it could name them.
A 1946 MIT housing lottery and a 2022 machine-learning study of the MLB draft, seventy years apart, found the identical bias: living close to a decision-maker predicts the outcome, even when it has nothing to do with merit.
Digital is cheaper to send, every time. Three real, cited studies on touch, reading comprehension, and how people value physical objects explain why organisations, NRMA's own quarterly member magazine among them, still mail the paper copy anyway.
Progress Flow (steps and a progress bar) and Progressive Load (one continuous page) both borrow real psychology, chunking and the goal-gradient effect, but six studies ranked by fidelity show the same pattern doing nothing on a desktop and backfiring on a phone.
Automation bias and algorithm aversion are the same miscalibration pointed in opposite directions, and keeping a human in the loop doesn’t reliably fix either one. Ten real, cited studies across AI coding assistants, clinical AI, and everyday knowledge work.
What a fair coin can teach you about statistics: why a real random sequence looks streakier than expected, why a run of heads doesn't make tails overdue, what a p-value actually says, and why the smallest sample in any ranking always looks the most extreme.
Ten million unintended duplicate super accounts, tens of thousands of members who sold at the exact bottom of the 2020 crash, and $37.8 billion pulled out early during the pandemic. Four real, cited mechanisms, and the one 2021 law change that turned the same default that caused the damage into the fix.
Internal, external, ecological, construct, and statistical conclusion validity: five separate questions a single result can pass or fail independently, each demonstrated on a real, cited replication case already documented on this page.
The outcome chain behind a feature launch, mapped stage by stage on one worked example: awareness, discovery, trial, onboarding, completion, and whether it actually persists, with the real proxy trap waiting at every link.
The decades-old evaluation discipline underneath every logic model: why a program fails for one of two very different reasons, shown on a real program that made crime worse while working exactly as designed, and one that got the theory right from day one.