Special Report

Every Program Runs on a Theory. Most Never Write It Down.

Program theory is a decades-old discipline in evaluation research, built on one distinction: a logic model draws the chain from activity to outcome, and a program theory explains why each link in that chain should actually hold. Most programs have the first. Very few write down the second.

This site's own report on outcome logic already built a version of this: five stages, one real metric per link, for a single savings-goal feature. That report is a logic model. This one is about the layer underneath it, the causal reasoning evaluators have spent fifty years formalising, and the two very different reasons a program actually fails.

One of the clearest illustrations on record is a program that was implemented almost exactly as designed, and made the thing it was trying to prevent worse. The theory was the problem, not the delivery. Getting that distinction right is what program theory is for.

The story in four parts

A program is a bet on a causal chain

Every initiative assumes that doing X will cause Y, whether or not anyone ever states the assumption in those words.

A program can fail for two different reasons

The underlying idea can be wrong, or a right idea can be delivered badly, and those two failures need completely different fixes.

A logic model draws the chain. Program theory explains it.

Naming the stages is the easy half. Stating why each transition should occur, and under what conditions, is the half that actually gets tested.

Writing the theory down is what makes it falsifiable

An unstated assumption can't be checked before the money's spent. A written one can.

The evidence at a glance

A well-delivered program increased the harm it was built to prevent

A Campbell Collaboration review of nine randomised and quasi-randomised trials found “Scared Straight”-style programs raised reoffending by 1 to 28% against untreated control groups.

506 villages, randomised, before a peso was spent nationally

Mexico's Progresa tested its own theory of change against 506 of 6,396 eligible villages before scaling, one of the earliest large randomised evaluations of a national anti-poverty programme.

Outcomes come in three tiers, not one

The W.K. Kellogg Foundation's own logic-model guide, already cited on this site, splits outcomes into short-term, medium-term, and long-term tiers, exactly the layered structure a bare goal statement skips.

A logic model draws the chain. Program theory explains why it should hold.

Evaluation researchers have kept these two things separate for decades. Huey-Tsyh Chen's foundational distinction splits a program's theory into a normative half (what the program does, and what it's meant to achieve) and a causative half (why doing that should actually produce the result), and argues evaluation needs both, not the normative half alone. Funnell and Rogers make the same split in practitioner terms: a logic model is the box-and-arrow diagram, and a full program theory adds the causal assumptions and the conditions under which each arrow actually holds.

This is the exact gap this site's own report on outcome logic runs into without naming it. Its five-stage funnel, with a real metric at every link, is a logic model. Program theory is the layer underneath: stating that a customer should move from “sent” to “opened” because attention has to cross a salience threshold, as an explicit, checkable claim, not an assumption nobody wrote down.

Sources: Chen, H. T. (1990). Theory-Driven Evaluations. Sage Publications. Funnell, S. C., & Rogers, P. J. (2011). Purposeful Program Theory: Effective Use of Theories of Change and Logic Models. Jossey-Bass. Both build on the same W.K. Kellogg Foundation (2004) Logic Model Development Guide already cited in Your Roadmap Has a Phase 2.

Once the theory is written down separately, a second question becomes possible

Naming the causal chain doesn't just describe a program, it makes a specific kind of diagnosis possible when the program fails. Carol Weiss's own framing of theory-based evaluation splits failure into two distinct causes: a theory failure, where the underlying causal idea itself was wrong, and an implementation failure, where a sound idea was delivered badly. Without an explicit program theory to check against, an evaluator watching a program fail cannot tell which one happened, and a wrong diagnosis sends the fix in the wrong direction entirely.

Source: Weiss, C. H. (1997). “Theory-Based Evaluation: Past, Present, and Future.” New Directions for Evaluation, 1997(76), 41–55.

Scared Straight is the sharpest real illustration of a theory failure

The distinction above is easy to state in the abstract. This is what a pure theory failure looks like in a real, repeatedly tested program. “Scared Straight”-style programs took juvenile offenders and at-risk youth into adult prisons to be confronted by inmates describing the reality of incarceration, on the causal theory that direct exposure to consequences would deter future crime.

Real Case

The real metric, and what it actually showed

Why did a program executed almost exactly as designed increase the crime it was built to prevent?

A Campbell Collaboration systematic review pooled nine randomised and quasi-randomised trials of Scared Straight and similar juvenile awareness programs, run across eight different U.S. states. Every study measured the same real outcome, subsequent offending, not attendance or satisfaction. The pooled result: reoffending rose 1 to 28% in the treatment group compared with untreated controls.

What it teaches: the programs were largely delivered as intended, prisoners spoke, tours ran, sessions happened. Better management would not have fixed this, because the causal assumption underneath it, that fear of a vivid future consequence reduces present behaviour, was itself wrong for this population. That's a theory failure, not an implementation failure, and only a stated theory makes the distinction checkable.

Petrosino, A., Turpin-Petrosino, C., Hollis-Peel, M. E., & Lavenberg, J. G. (2013). “Scared Straight and Other Juvenile Awareness Programs for Preventing Juvenile Delinquency: A Systematic Review.” Campbell Systematic Reviews, 9(1), 1–55

Progresa is the reverse case: a theory specific enough to be tested

Scared Straight shows what happens when the theory is wrong and nobody checked before scaling it nationally. Mexico's Progresa programme shows the other direction. Designed in 1997 by deputy finance minister Santiago Levy, Progresa (later renamed Oportunidades) paid cash transfers to mothers in poor households, conditional on children attending school and family members attending health checkups.

Real Case

The real metric, and what it actually showed

What made Progresa's theory testable instead of just hoped for?

The theory was stated as a specific causal chain, not a vague goal: cash conditioned on school attendance and health checkups would increase human-capital investment, which would reduce poverty in the next generation, not just relieve it in the current one. That specificity is what let evaluators randomise 506 of the 6,396 eligible villages into treatment and control groups before the programme rolled out nationally, one of the earliest large-scale randomised evaluations of a national anti-poverty programme anywhere.

What it teaches: a well-specified theory doesn't guarantee a program works. It guarantees that whether it works, and why, can actually be tested, which is the whole point of writing one down in the first place.

Read together, these two cases teach the same lesson from opposite directions

Scared Straight and Progresa are not being compared because one is good and one is bad. Both are evidence for the same underlying claim. A named, specific program theory is what made Scared Straight's failure diagnosable as a theory failure rather than a shrug, and it's what let Progresa's designers test their actual causal claim before betting a national budget on it. Neither outcome was visible in advance. The theory being written down and specific enough to check is what made each outcome legible after the fact.

Before you fund a program, or fund it again, four checks

This is the practitioner's version of the discipline above. Four checks, not a full evaluation plan, because a program theory that never gets written down because the process felt too heavy defeats the entire point.

  1. Have you written the causal chain down, not just the goal?“Reduce poverty” and “reduce delinquency” are goals. The chain is the specific steps that are supposed to produce them.
  2. For each link, can you state why it should hold, and under what conditions?“Kids will be scared into compliance” is a claim. Stating why, and for which kids, is the theory.
  3. If this program fails, could you tell whether the theory was wrong or the delivery was?Without a written theory to check against, a failed program only ever tells you it failed, not why.
  4. Has anyone tried to break the theory on paper, before spending money testing it in the field?Progresa's theory survived that scrutiny before scaling. Scared Straight's theory never got the same test.