A study can be internally airtight, statistically significant, and still tell you almost nothing useful. That's not a contradiction. Passing one methodological test says nothing about whether it passed the other four, and a sentence like “the study found X” almost never says which ones it actually cleared.
This site already leans on two of these words a lot: internal validity and external validity, both used throughout Reading the Research. They aren't the only two that matter. Conflating them with the other three, construct validity, statistical conclusion validity, and ecological validity, is exactly how a real, honestly-run study ends up applied to the wrong question.
This report walks through all five, each one demonstrated on a real, cited case already documented on this site. One question ties them together: what does this specific result actually license you to believe, and what does it not?
The story in four parts
An unblinded measurement can produce a genuine result that has nothing to do with the treatment being tested.
A finding can hold up perfectly in the exact setting it was tested in, and vanish the moment the setting changes even slightly.
A metric can move in a completely real, statistically solid way, while standing in for a concept it never actually captured.
Years of small, individually plausible studies can add up to a consensus that a single, adequately powered test dissolves.
The evidence at a glance
A 2012 replication of a famous priming study found no effect with objective timing, and an effect only when the humans running the stopwatch believed one would appear.
A well-known facial-feedback study replicated across 17 independent labs vanished the moment participants knew they were being recorded on video.
A real bank's own data found daily-login engagement negatively correlated with financial wellbeing, the opposite of what the metric was assumed to represent.
Decades of studies said willpower depletes like a muscle. A single adequately powered, pre-registered replication across 23 labs found no meaningful effect at all.
Internal validity asks the narrowest and most fundamental question of the five: within this exact study, did the treatment actually cause the effect measured, or could something else explain it just as well? Confounding variables, selection effects, and the study's own measurement process are the usual suspects. That last one is easy to overlook, because it feels like part of collecting the data rather than part of the experiment.
Why did people only walk slower after old-age words when the experimenter expected it?
A hugely influential study found participants exposed to words associated with old age walked measurably slower leaving the lab afterwards. A 2012 replication using objective infrared timing, rather than a human holding a stopwatch, found no priming effect at all. The effect only reappeared in a second experiment, when the person running the stopwatch believed in advance that participants would walk slower.
What it teaches: if whoever measures the outcome knows or can guess the hypothesis and isn't blinded to it, their own expectations can leak into a subjective judgement, “how slowly did that person seem to walk.” Blinded, objective measurement is specifically what internal validity protects against.
External validity asks whether a result generalises: to a different population, a different setting, a different time, would the same effect still show up? A study can be flawlessly internally valid. The treatment really did cause the outcome in that room. It can still describe a narrow, specific world that doesn't match the one a decision is actually being made in.
Why did the pen-in-the-teeth effect disappear the moment anyone was watching on video?
The original study found people rated cartoons funnier when a pen held in their teeth forced their face into a smile. A 2016 replication across 17 labs and over 2,000 participants, following a protocol the original author helped design, found no effect at all. One proposed explanation: the replication labs, unlike the original, recorded participants on video, and being watched appears to suppress the effect.
What it teaches: a failed replication doesn't always mean the original effect was fake. Sometimes it means the effect has a real boundary condition, here, being observed, that the original study never tested for and the finding quietly depended on all along.
Ecological validity is easy to mistake for external validity, and it's worth keeping the two separate. External validity asks whether a result generalises to other people or settings. Ecological validity asks a narrower question: does the task inside the study itself resemble the real behaviour it claims to measure? A perfectly internally valid, perfectly generalisable finding about a task can still fail this test, if the task was never a realistic stand-in for what people actually do.
Why do people say they'd pay more than they actually do, even when warned about it?
Surveys asking “would you pay for this?” measure something real, but not the thing they're usually assumed to measure. When there's no real cost to answering however makes you look good, people reliably state a higher willingness to pay in a hypothetical survey than they actually pay when real money is on the line. Warning respondents about this bias beforehand has generally proven weak at closing the gap.
What it teaches: a stated intention answers a real question, just not the one it's usually taken to answer. It's a lower-stakes task standing in for a higher-stakes one, and the two don't reliably move together.
Construct validity asks whether the number actually being measured represents the underlying idea it's assumed to stand in for. “Engagement,” “loyalty,” and “satisfaction” are all constructs, not directly observable things, and the metric chosen to operationalise one of them can be completely real, statistically sound, and still measuring something else entirely.
Why were a bank's most engaged customers its most financially stressed ones?
A real study, combining a nationally representative US survey with administrative data from 194,678 accounts at a large Australian bank, found daily app engagement was negatively correlated with financial wellbeing (r = -0.34). The most digitally engaged customers were the least financially well, the opposite direction the common “engagement equals loyalty” assumption predicts.
What it teaches: the correlation is completely real. The problem sits one level up, in what “engagement” was assumed to represent. A customer checking their balance daily out of financial anxiety produces the identical metric as one checking it daily out of genuine satisfaction, and the metric alone can't tell those two people apart.
This is the question the other four don't ask. Even with a perfect design, a real population, and a metric that genuinely captures the intended concept, an underpowered test can manufacture an effect that isn't there, or miss a real one entirely. This site's own coin-flip report covers the mechanics of this in full, with exact numbers. The case below shows what it looks like when it happens to a famous, real finding instead of a hypothetical coin.
Why did a finding from 42 people become one of the most famous results in behavioural science?
A 2010 study of 42 participants found holding an expansive “power pose” for two minutes raised testosterone, lowered cortisol, and increased self-reported feelings of power. It became one of the field's most famous findings, boosted by a TED talk viewed over 20 million times. A 2015 replication with a properly powered sample of 200 reproduced the self-reported feeling of power, but found no hormonal or behavioural effect at all.
What it teaches: a small sample can produce a result that looks completely clean and still be noise. One of the original co-authors has since said publicly she no longer believes the hormonal and behavioural effects are real.
None of these five failures require dishonesty, and most of the studies above were run by careful, well-intentioned researchers. They're default ways a result can mislead, which is exactly why they're worth checking for deliberately, one at a time.