← Studies Suggest 🧠 Psychology

The Marshmallow Test Convinced the World That a Four-Year-Old's Willpower Predicts Adult Success. A 918-Child Replication Cut the Effect by Two Thirds, and Psychologists Are Still Fighting Over the Rest

Watts, Duncan, and Quan re-ran the most famous experiment in developmental psychology with a sample ten times larger than the original and found the celebrated link between waiting for a second marshmallow and later achievement was half the original size, then shrank by two thirds once family background entered the model. Two published rebuttals argue they controlled away the very thing the test measures.

By Rebecca Liang, Behavioral Psychology - September 8, 2026

Listen to this article
Loading...
A single white marshmallow resting on a weathered wooden table in warm morning light, soft-focus green foliage behind, shallow depth of field, earthy greens and warm cream tones, nature magazine style

πŸ“‹ The Study

TitleRevisiting the Marshmallow Test: A Conceptual Replication Investigating Links Between Early Delay of Gratification and Later Outcomes
AuthorsWatts, Duncan & Quan, 2018
InstitutionNew York University (Watts); University of California, Irvine (Duncan, Quan)
JournalPsychological Science, 29(7), 1159–1177
DOI10.1177/0956797618761661
Samplen=918 children (NICHD Study of Early Child Care and Youth Development); focal analysis n=552 children whose mothers had not completed college
MethodLongitudinal conceptual replication: delay of gratification measured at 54 months, academic and behavioral outcomes at age 15, with covariate-adjusted regression models
Key FindingEach extra minute waited at age 4 predicted ~0.10 SD more achievement at age 15, but the raw link was half the original size and shrank by two thirds after controlling for family background, early cognition, and home environment
Effect SizeBivariate: ~0.10 SD achievement per extra minute waited. With background controls: roughly one third of that. With concurrent cognitive/behavioral controls: nearly zero. Original 1990 SAT correlations: r=.57 (math) and r=.42 (verbal), n=37.
Counterintuition⚑⚑⚑⚑ 4/5
ReplicationChallenged and actively debated - conceptual replication whose interpretation was contested by two published critiques (Falk, Kosse & Pinger, 2020; Doebel, Michaelson & Munakata, 2020), with an author reply (Watts & Duncan, 2020). No retraction; no corrections per Crossref; no PubPeer flags found; 306+ citations.

The Most Famous Four-Year-Olds in Science

In the early 1970s, psychologist Walter Mischel began offering Stanford preschoolers a stark bargain: one marshmallow now, or two if you can sit alone with the treat for fifteen minutes. The footage became iconic: small children squirming, sniffing the treat, covering their eyes, bargaining with an empty room. Some last the full quarter hour, while others last eleven seconds.

Years later, Mischel's team tracked the children down and reported something astonishing. In the 1990 follow-up, the preschoolers who had waited longest scored dramatically higher on the SAT, with correlations of .57 for math and .42 for verbal. Those numbers, computed from just 37 children, became some of the most quoted statistics in psychology, and the moral hardened into conventional wisdom for three decades of parenting books and TED talks. Self-control at four predicts success for life, a thesis Mischel's 2014 book wore in its subtitle. It did not. Teach your child to wait, and the future takes care of itself.

A Bigger, Broader Rerun

In 2018, Tyler Watts, Greg Duncan, and Haonan Quan checked the math with a sample the original never had. They drew on the NICHD Study of Early Child Care and Youth Development, which had followed children from birth at ten US sites, assembling 918 children with a delay measure at 54 months and outcome data at age 15. Most analyses focused on 552 children whose mothers had not completed college: ten times the original longitudinal sample and far more representative of American families.

The headline result looked like confirmation, then did not. Each extra minute waited at age four predicted about a tenth of a standard deviation more achievement at fifteen, but the authors stressed it was only half the size of the originals. It shrank by two thirds after accounting for family background, early cognitive ability, and the home environment, and models that also controlled for concurrent cognitive and behavioral skills reduced it to nearly zero, which is the statistical equivalent of watching the most famous effect in developmental psychology dissolve. Links to age-15 behavior were smaller still and rarely significant.

Then came the detail that reframes the test. Nearly all the predictive signal came from one threshold: waiting at least twenty seconds. The gap between instant grabbing and a half-minute hold carried the association; the heroic quarter-hour performances immortalized on video added almost nothing. The famous footage shows the least informative part of the experiment.

The Incredible Shrinking Correlation

A calculation the paper never ran makes the shrinkage visceral. Take the headline .57 for SAT math: squared, it says waiting explains about 32 percent of variance. Halve the correlation, as the replication's raw result roughly does, and the explained variance falls to 8 percent; cut it by two thirds more and you land under 1 percent. From nearly a third of the variance to less than one percent: a thirty-six-fold collapse in explanatory power. The arithmetic applies the paper's own "half, then a third" to the headline correlation, so read it as illustration; the order of magnitude should send anyone citing the original .57 reaching for a confidence interval first.

A second humbling number: with 37 children behind the SAT-math correlation, the 95 percent confidence interval stretches from roughly .30 to .75. The most celebrated number in the literature was never precise: a blurry estimate from a tiny sample, with nearly the entire replication debate fitting inside its original error bars.

The Rebuttals Deserve Full Weight

The critics did not accept the shrinkage quietly.

Economists Armin Falk, Fabian Kosse, and Pia Pinger reanalyzed both datasets in 2020 and concluded the new data confirm the original finding: the correlation between willingness to delay and later school success is relatively strong in the replication too, which suggests the debunking says more about covariates than willpower. On this reading, the replication did not debunk the test; it showed the effect in a broader population than anyone had demonstrated. The interesting question is what determines whether a child can wait, not whether waiting matters.

Sabine Doebel, Laura Michaelson, and Yuko Munakata attacked the controls themselves. Many covariates, they argued, measure the very machinery of waiting: executive function, verbal ability, home environment, and trust that rewards arrive. Controlling for the foundations of self-control and then announcing that self-control predicts nothing is circular, like controlling for engines and concluding cars need none; strip the over-control, and the results read as a successful partial replication, not a refutation.

Watts and Duncan replied in 2020, defending the line between confounds and the construct itself; the exchange remains unresolved.

What We Didn't Prove

This was a conceptual replication, not a rerun: different decades, a broader and less affluent sample, a modified task, and an NICHD sample that was geographically diverse but not nationally representative. The two studies are not perfectly comparable, which is what a clean verdict would require.

A deeper limitation cuts both ways. Every number here is correlational: if waiting predicts achievement because both reflect a stable home, that is the challengers' story; if waiting expresses a deeper capacity the controls erased, that is the defenders'. The data cannot adjudicate between a confound and a mechanism when they are the same variables, which means the whole debate hinges on a modeling choice no dataset can make for you. Neither side has run the decisive experiment, which would randomly boost some children's ability to wait and nothing else, then check who thrives at fifteen.

One more caveat favors the test's skeptics. Celeste Kidd's 2013 experiments showed children wait less when the adult running the test proves unreliable. A four-year-old who eats the marshmallow may have rationally judged that the second marshmallow is never coming. The test measures trust in the room as well as discipline in the child.

The Bottom Line

The marshmallow test is not dead, but the simple parable is over. The strongest version of the original claim, that a preschooler's waiting time explains later achievement on its own, does not survive 918 children and a few background variables. What survives is smaller and more interesting: the ability to wait briefly, nearly all of it in the first twenty seconds, marks already-advantaged children. Whether that counts as replication or refutation depends on which covariates belong in the model, which is why the field still argues, and why both sides can look at the same 918 children and see vindication, which should embarrass anyone treating one regression as the voice of nature. Science at its best is not a gavel; it is a ledger that stays open.

What You Can Do

First, retire the marshmallow test as a parenting diagnostic: a four-year-old who eats the treat is telling you about that afternoon, not her future.

Second, invest where the controls point. Reading, conversation, routines, and play build what the test only gestures at.

Third, be someone whose promises hold. Kidd's reliability finding is the most directly actionable here: children calibrate waiting to whether adults deliver, so keep your word about the second marshmallow, literally and figuratively.

Fourth, target broad capacities, not the drill: programs building executive function across contexts have a plausible mechanism, while programs that merely train children to stare down candy do not.

Sources

  1. Watts, T. W., Duncan, G. J., & Quan, H. (2018). Revisiting the marshmallow test: A conceptual replication investigating links between early delay of gratification and later outcomes. Psychological Science, 29(7), 1159–1177. https://doi.org/10.1177/0956797618761661
  2. Shoda, Y., Mischel, W., & Peake, P. K. (1990). Predicting adolescent cognitive and self-regulatory competencies from preschool delay of gratification: Identifying diagnostic conditions. Developmental Psychology, 26(6), 978–986. https://doi.org/10.1037/0012-1649.26.6.978
  3. Mischel, W., Shoda, Y., & Rodriguez, M. L. (1989). Delay of gratification in children. Science, 244, 933–938. https://doi.org/10.1126/science.2658056
  4. Falk, A., Kosse, F., & Pinger, P. (2020). Re-revisiting the marshmallow test: A direct comparison of studies by Shoda, Mischel, and Peake (1990) and Watts, Duncan, and Quan (2018). Psychological Science, 31(1), 100–104. https://doi.org/10.1177/0956797619861720
  5. Watts, T. W., & Duncan, G. J. (2020). Controlling, confounding, and construct clarity: Responding to criticisms of "Revisiting the marshmallow test" by Doebel, Michaelson, and Munakata (2020) and Falk, Kosse, and Pinger (2020). Psychological Science, 31, 105–108. https://doi.org/10.1177/0956797619893606
  6. Doebel, S., Michaelson, L. E., & Munakata, Y. (2020). Good things come to those who wait: Delaying gratification likely does matter for later achievement (a commentary on Watts, Duncan, & Quan, 2018). Psychological Science. https://pmc.ncbi.nlm.nih.gov/articles/PMC10152116/
  7. Kidd, C., Palmeri, H., & Aslin, R. N. (2013). Rational snacking: Young children's decision-making on the marshmallow task is moderated by beliefs about environmental reliability. Cognition, 126(1), 109–114. https://doi.org/10.1016/j.cognition.2012.08.004
  8. Casey, B. J., Somerville, L. H., Gotlib, I. H., et al. (2011). Behavioral and neural correlates of delay of gratification 40 years later. Proceedings of the National Academy of Sciences, 108(36), 14998–15003. https://doi.org/10.1073/pnas.1108561108