How do you prove distance biased 30,000 draft picks when nobody can randomly assign where a scout lives?
Major League Baseball's amateur draft is a real, ongoing natural experiment already running at scale: 30 scouting directors, evaluating thousands of prospects a year, none of them assigned to their job or their home city by a coin flip. A 2022 study used that fact to ask a pointed question. Once you strip out actual skill, does living closer to a team's scouting director make a player more likely to get drafted by that team, and does it cost the team anything if it does?
You can't answer that with a randomised trial. Nobody can randomly assign where a scouting director lives, or randomly move a prospect's hometown. This entry breaks down what the researchers did instead, and why it holds up as real evidence despite never touching a control group.
The honest question here is causal: does proximity itself change a scouting director's decision, or do nearby players just happen to be better, because a director naturally scouts their own backyard more thoroughly and finds real talent there? A randomised trial would settle this outright, randomly assign scouting directors to team markets, or randomly assign prospects to hometowns, and see if the draft outcome moves.
Neither is remotely possible. Scouting directors are hired for reasons that have nothing to do with a researcher's study design, and a player's hometown is fixed before anyone involved has ever heard of them. Whatever answer this gets has to come from data the draft already generated on its own, not from a trial anyone could run.
The 2022 study, covering the 2000–2019 MLB drafts, roughly 30,000 players drafted from a pool of over a million, used machine learning to build a flexible prediction of each prospect's "true" draft-worthy skill from everything else on record about them: their statistics, their scouting grades, their physical measurables, all the information a real scouting department would have had.
With that skill estimate held constant, the remaining question becomes narrow and answerable: among two prospects the model rates as equally good, does the one living closer to the scouting director still get picked more often? The machine learning model is doing the same job random assignment would have done in a trial, removing the confound (real skill) that would otherwise explain the whole pattern, just doing it statistically after the fact instead of physically before the study began.
A single correlation surviving one set of controls wouldn't be enough on its own, and the study doesn't rest on just one. The bias shows up specifically in the draft's later rounds, exactly where a scouting director has the most personal discretion and the least outside scrutiny, and is measurably weaker in the early rounds, where a consensus national ranking constrains any one director's opinion.
That pattern is hard to explain as a leftover statistical artifact: an artifact of bad controls should show up everywhere skill is hard to measure, not specifically concentrate itself in the exact organizational moment where one person's personal judgment carries the most weight. A finding that gets stronger exactly where the proposed mechanism predicts it should, and weaker exactly where it predicts it shouldn't, is doing real evidentiary work that a single flat effect size wouldn't.
A player is roughly 7.1% more likely to be drafted by a given team for every 1,000km he lives closer to that team's scouting director, and about 4.9% more likely for every 1,000km closer to the team's home city, holding the model's skill estimate constant.
It isn't a harmless quirk: players drafted with that proximity advantage were 38% less likely to ever appear in a single MLB game than a similarly-drafted player without it, and appeared in roughly 25 fewer games on average. The bias didn't just reshuffle who got picked when. It measurably lowered the quality of who got picked at all.
- A real, consistent association between geographic proximity and draft outcome that survives controlling for every measurable indicator of skill available in the data.
- A pattern (strongest where scouts have the most discretion) that matches what a genuine bias would predict, not just a flat, unexplained correlation.
- A real cost: proximity-favoured picks perform measurably worse, so this isn't a victimless statistical curiosity.
- Some skill that exists in the real world but wasn't captured by any of the model's inputs could still be doing part of the work, no observational control can promise it caught every confound a true random assignment would have eliminated automatically.
- The paper is still a working paper, not yet a peer-reviewed publication at time of writing, worth checking before treating its exact numbers as final.
- It's one league, one two-decade window. Whether the same size effect holds in a different sport, country, or era hasn't been separately tested here.
Ahmadi, M., Durst, N., Lachman, J., List, J. A., List, M., List, N., & Vayalinkal, A. (2022). “Nothing Propinks Like Propinquity: Using Machine Learning to Estimate the Effects of Spatial Proximity in the Major League Baseball Draft.” NBER Working Paper No. 30786
Every number above is paraphrased from the paper's own public summary and independent academic coverage, not quoted from the full text, which sat behind this site's network restrictions.
See also: Propinquity Effect, the underlying psychological mechanism this draft study demonstrates at professional scale, first documented in a very different setting, a 1950s housing complex. How Strong Is That Evidence?, this site's own field guide to where a study like this sits on the ladder from raw correlation to a randomised trial. And Proximity Gets Mistaken for Merit, a special report reading the housing study and this draft study together as one argument.