The Most Famous Four-Year-Olds in Science
In the early 1970s, psychologist Walter Mischel began offering Stanford preschoolers a stark bargain: one marshmallow now, or two if you can sit alone with the treat for fifteen minutes. The footage became iconic: small children squirming, sniffing the treat, covering their eyes, bargaining with an empty room. Some last the full quarter hour, while others last eleven seconds.
Years later, Mischel's team tracked the children down and reported something astonishing. In the 1990 follow-up, the preschoolers who had waited longest scored dramatically higher on the SAT, with correlations of .57 for math and .42 for verbal. Those numbers, computed from just 37 children, became some of the most quoted statistics in psychology, and the moral hardened into conventional wisdom for three decades of parenting books and TED talks. Self-control at four predicts success for life, a thesis Mischel's 2014 book wore in its subtitle. It did not. Teach your child to wait, and the future takes care of itself.
A Bigger, Broader Rerun
In 2018, Tyler Watts, Greg Duncan, and Haonan Quan checked the math with a sample the original never had. They drew on the NICHD Study of Early Child Care and Youth Development, which had followed children from birth at ten US sites, assembling 918 children with a delay measure at 54 months and outcome data at age 15. Most analyses focused on 552 children whose mothers had not completed college: ten times the original longitudinal sample and far more representative of American families.
The headline result looked like confirmation, then did not. Each extra minute waited at age four predicted about a tenth of a standard deviation more achievement at fifteen, but the authors stressed it was only half the size of the originals. It shrank by two thirds after accounting for family background, early cognitive ability, and the home environment, and models that also controlled for concurrent cognitive and behavioral skills reduced it to nearly zero, which is the statistical equivalent of watching the most famous effect in developmental psychology dissolve. Links to age-15 behavior were smaller still and rarely significant.
Then came the detail that reframes the test. Nearly all the predictive signal came from one threshold: waiting at least twenty seconds. The gap between instant grabbing and a half-minute hold carried the association; the heroic quarter-hour performances immortalized on video added almost nothing. The famous footage shows the least informative part of the experiment.
The Incredible Shrinking Correlation
A calculation the paper never ran makes the shrinkage visceral. Take the headline .57 for SAT math: squared, it says waiting explains about 32 percent of variance. Halve the correlation, as the replication's raw result roughly does, and the explained variance falls to 8 percent; cut it by two thirds more and you land under 1 percent. From nearly a third of the variance to less than one percent: a thirty-six-fold collapse in explanatory power. The arithmetic applies the paper's own "half, then a third" to the headline correlation, so read it as illustration; the order of magnitude should send anyone citing the original .57 reaching for a confidence interval first.
A second humbling number: with 37 children behind the SAT-math correlation, the 95 percent confidence interval stretches from roughly .30 to .75. The most celebrated number in the literature was never precise: a blurry estimate from a tiny sample, with nearly the entire replication debate fitting inside its original error bars.
The Rebuttals Deserve Full Weight
The critics did not accept the shrinkage quietly.
Economists Armin Falk, Fabian Kosse, and Pia Pinger reanalyzed both datasets in 2020 and concluded the new data confirm the original finding: the correlation between willingness to delay and later school success is relatively strong in the replication too, which suggests the debunking says more about covariates than willpower. On this reading, the replication did not debunk the test; it showed the effect in a broader population than anyone had demonstrated. The interesting question is what determines whether a child can wait, not whether waiting matters.
Sabine Doebel, Laura Michaelson, and Yuko Munakata attacked the controls themselves. Many covariates, they argued, measure the very machinery of waiting: executive function, verbal ability, home environment, and trust that rewards arrive. Controlling for the foundations of self-control and then announcing that self-control predicts nothing is circular, like controlling for engines and concluding cars need none; strip the over-control, and the results read as a successful partial replication, not a refutation.
Watts and Duncan replied in 2020, defending the line between confounds and the construct itself; the exchange remains unresolved.
What We Didn't Prove
This was a conceptual replication, not a rerun: different decades, a broader and less affluent sample, a modified task, and an NICHD sample that was geographically diverse but not nationally representative. The two studies are not perfectly comparable, which is what a clean verdict would require.
A deeper limitation cuts both ways. Every number here is correlational: if waiting predicts achievement because both reflect a stable home, that is the challengers' story; if waiting expresses a deeper capacity the controls erased, that is the defenders'. The data cannot adjudicate between a confound and a mechanism when they are the same variables, which means the whole debate hinges on a modeling choice no dataset can make for you. Neither side has run the decisive experiment, which would randomly boost some children's ability to wait and nothing else, then check who thrives at fifteen.
One more caveat favors the test's skeptics. Celeste Kidd's 2013 experiments showed children wait less when the adult running the test proves unreliable. A four-year-old who eats the marshmallow may have rationally judged that the second marshmallow is never coming. The test measures trust in the room as well as discipline in the child.
The Bottom Line
The marshmallow test is not dead, but the simple parable is over. The strongest version of the original claim, that a preschooler's waiting time explains later achievement on its own, does not survive 918 children and a few background variables. What survives is smaller and more interesting: the ability to wait briefly, nearly all of it in the first twenty seconds, marks already-advantaged children. Whether that counts as replication or refutation depends on which covariates belong in the model, which is why the field still argues, and why both sides can look at the same 918 children and see vindication, which should embarrass anyone treating one regression as the voice of nature. Science at its best is not a gavel; it is a ledger that stays open.
What You Can Do
First, retire the marshmallow test as a parenting diagnostic: a four-year-old who eats the treat is telling you about that afternoon, not her future.
Second, invest where the controls point. Reading, conversation, routines, and play build what the test only gestures at.
Third, be someone whose promises hold. Kidd's reliability finding is the most directly actionable here: children calibrate waiting to whether adults deliver, so keep your word about the second marshmallow, literally and figuratively.
Fourth, target broad capacities, not the drill: programs building executive function across contexts have a plausible mechanism, while programs that merely train children to stare down candy do not.