💰 Economics
Nudges Don’t Scale. 126 Government Trials Covering 23 Million People Found Behavioral Tweaks Work at One-Sixth the Strength the Journals Claim
Academic nudge studies report an 8.7-point lift in take-up, a 33% jump. The same kind of interventions run by government nudge units managed 1.4 points. The gap is mostly publication bias.
- Title
- RCTs to Scale: Comprehensive Evidence from Two Nudge Units
- Authors
- DellaVigna S, Linos E (2022)
- Institution
- University of California, Berkeley (and NBER)
- Journal
- Econometrica 90(1):81-116
- Sample
- 126 RCTs, 243 nudge treatments, 23M+ participants (two US nudge units); 26 papers, 74 treatments, ~505K participants (academic journals). ⚠️ US government settings, 2015-2019; binary take-up outcomes only; defaults and financial incentives excluded.
- Method
- Complete trial census from two nudge units (Office of Evaluation Sciences + Behavioural Insights Team North America) compared with published academic nudge RCTs from two meta-analyses under identical restrictions; meta-analytic model decomposing the effect-size gap into power, intervention characteristics, and selective publication
- Key Finding
- Academic nudges averaged +8.7pp take-up (+33.4%); nudge-unit nudges averaged +1.4pp (+8.0%), about one sixth the size, still highly significant
- Effect Size
- 8.7pp vs 1.4pp (6.2x ratio); correcting for publication bias pulls the academic estimate to 3.2pp; null results estimated 10x less likely to be published
- Counterintuition
- ⚡⚡⚡⚡ 4/5
- Replication
- Not yet replicated: first comprehensive census of its kind, so there was no sampling and no scope for selective publication within it; replication files are public via the Econometric Society. The academic comparison leans on two published meta-analyses (Benartzi et al. 2017; Hummel & Maedche 2019).
In 2008, Richard Thaler and Cass Sunstein's Nudge gave policymakers a seductive idea: you do not need to ban anything, tax anything, or spend much of anything, just redesign the choice architecture. Simplify the form, personalize the letter, tell people what their neighbors do. Tiny tweaks, they argued, could move behavior predictably without forbidding options or significantly changing economic incentives. Governments listened. By one OECD count, more than 200 behavioral science units now operate inside governments worldwide, running experiments on everything from tax collection to retirement savings.
Two Berkeley economists decided to check what happens when the hits leave the lab. Stefano DellaVigna and Elizabeth Linos obtained something researchers almost never get, the complete files of two of the largest nudge units in the United States: first, the Office of Evaluation Sciences, launched in 2015 as the core of the Obama White House's Social and Behavioral Sciences Team and created by executive order to embed behavioral science across federal agencies. The second was the Behavioural Insights Team's North American office, which spent the period working with more than fifty American cities. Both units had kept a comprehensive record of every trial since their 2015 inception, each with a trial report and often a pre-analysis plan. As of July 2019 they had run 165 trials testing 349 nudge treatments on more than 37 million participants, and 87 percent had never appeared in any working paper or journal. There was no file drawer here. The researchers got the whole filing cabinet.
The analysis sample came down to 126 randomized trials, 243 distinct nudges, and more than 23 million people. The comparison sample came from the academic literature: 26 published papers containing 74 nudge treatments and half a million participants, drawn from two recent meta-analyses under the same restrictions. Financial incentives and defaults were excluded, and every outcome was a binary take-up measure, so the metric was always the same: percentage points of additional take-up versus control. The interventions were the nudge toolkit: simplification, personalization, reminders, implementation-intention prompts, and social-norm comparisons delivered through administrative mail. One trial sent servicemembers a letter encouraging them to reenroll in their Roth Thrift Savings Plan; another sent residents a postcard urging them to fix up their homes to meet code regulations.
The journals said nudges were transformative. The filing cabinet said otherwise. Across the 26 academic papers, the average nudge lifted take-up by 8.7 percentage points, a 33.4 percent increase over the 26.0 percent baseline. Across the 126 nudge-unit trials, the average lift was 1.4 percentage points, an 8.0 percent increase over a 17.2 percent baseline. The authors summarized the at-scale effect as about one sixth the size of the literature's, while stressing that it remained highly statistically significant and, in their words, sizable, since a 1.4-point lift at roughly zero marginal cost is a return on investment most programs would envy.
The paper quantified the explanation in three parts. First came statistical power: the median academic treatment arm held 484 participants, enough to reliably detect only effects of 6.3 percentage points or larger, while the median nudge-unit arm held 10,006 participants, sensitive down to 0.8 points, a full order of magnitude of difference in what each literature was even capable of seeing. Second, the interventions genuinely differed: journal studies used more in-person contact and more choice design, worked different policy areas, and faced fewer institutional constraints. And third came selective publication: in the academic sample, there were more than four times as many results with t-statistics just above the significance threshold as just below it, the fingerprint of a literature that buries its nulls. The authors stress that “selective publication” includes studies researchers never bothered to write up, not just papers journals rejected. Their model estimated that null results were ten times less likely to see daylight than significant ones. Correcting for the censoring pulled the academic estimate down to 3.2 percentage points, still above the nudge-unit figure but no longer significantly so. The decomposition: publication bias, worsened by low power, explained about 70 percent of the gap, with intervention differences covering most of the rest.
The authors surveyed academic researchers and nudge practitioners, asking each group to forecast the results, and the split was stark: the median practitioner predicted a 1.95-point lift, close to the actual 1.4, while the median academic predicted 4.0 points for the nudge-unit trials and 7.0 for the journal studies. The people running the trials knew. The literature was the thing that was confused, and the confusion was built into its incentives: if you expect large effects, you run small trials, and small trials plus a file drawer manufacture large published effects.
Now the article's own calculation, with the inputs shown. Divide the journals' 8.7 points by the field's 1.4 and you get 6.2: the published number is more than six times the at-scale number. In relative terms, 33.4 percent against 8.0 percent, a policymaker budgeting on the literature overestimates the lift by a factor of 4.2. Make it concrete. A city mails a simplified reenrollment letter to 100,000 residents, expecting the journals' 8.7-point lift to produce 8,700 new savers, while the at-scale evidence says to expect about 1,400. That is still 1,400 savers conjured from stationery, which is the honest version of the nudge promise; the dishonest version is the other 7,300.
Now the strongest case against this reading, at full strength. The two samples were never identical twins: the nudge units operated under institutional constraints that largely ruled out default changes, the heavy artillery of choice architecture that the literature shows has the largest effects, leaving mostly one-shot letters and postcards, while the academic studies enjoyed more in-person contact and richer choice design. The authors themselves call the 1.4-point estimate a lower bound on what behavioral science can do at scale. There is also a fair complaint about where the “bias” lives, since the practitioners were nearly perfectly calibrated, which means the nudge units were never fooling themselves; the distortion sits in the academic journals, in the file drawer, in the incentive to publish the significant. And the most important caveat is the one the paper states plainly: 1.4 points at zero marginal cost is still sizable. Nobody in this literature, including its fiercest critics, claims nudges do nothing. The claim under test was ever the magnitude, and magnitude is where the hype lived.
That distinction rescues the idea while burying the sales pitch. Thaler and Sunstein never promised that a postcard could replace a tax. Their definition of a nudge explicitly excludes significantly changing economic incentives, which means the strongest interventions were outside the concept from the start. What the 126 trials actually show is something more useful than the legend, that light-touch behavioral interventions reliably move behavior a little, at almost no cost, across tens of millions of people, even though the published literature systematically exaggerated how much, which means a policymaker who budgets for 1.4 points and gets them has a working tool. A policymaker who budgets for 8.7 has a disappointment manufactured in a file drawer.
Small, cheap, and real is the fair description of the nudge at scale; the journals sold transformative, the filing cabinet says one sixth, so believe the cabinet.
What We Didn't Prove
- These are US government settings from 2015 to 2019. Results may not generalize to other countries, time periods, or non-government contexts.
- Only binary take-up outcomes were analyzed. Default changes were excluded by design, as were trials using financial incentives.
- The academic comparison sample is 26 papers drawn from two meta-analyses: representative of the literature, not exhaustive of it.
- The 70% attribution to publication bias comes from a meta-analytic model with its own assumptions, not a direct observation of unpublished studies.
- Averages are unweighted across heterogeneous interventions; individual nudges vary widely around the 1.4-point mean.
- The authors describe 1.4 points as a lower bound for behavioral science at scale, given the institutional constraints on the nudge units.
The Bottom Line
Across 126 government randomized trials covering 23 million people, nudges lifted take-up by 1.4 percentage points, about one sixth the 8.7-point average reported in academic journals. Publication bias, worsened by underpowered trials, explains roughly 70% of the gap. Nudges work. The literature exaggerated how much, by a factor of six.
What You Can Do
- If you make policy, budget on 1 to 3 points, not 8. Demand pre-registered, adequately powered trials before scaling: the nudge units’ median arm held 10,006 people, against 484 in the journals. And publish your nulls.
- If you run experiments, power them. A 484-person arm can only reliably detect effects above 6.3 points. Small trials plus a file drawer manufacture large published effects; the file drawer starts at your desk, not the journal’s.
- If you read science news, ask for the denominator. Twenty-six published papers versus 165 trials is the whole story in one ratio.
- Do not discard the tool. A 1.4-point lift at roughly zero marginal cost, applied to millions of people, is real behavior change. Pair light-touch nudges with stronger instruments, like defaults, where institutions allow them.
Sources
- DellaVigna S, Linos E. RCTs to Scale: Comprehensive Evidence from Two Nudge Units. Econometrica. 2022;90(1):81-116. doi:10.3982/ECTA18709 (working paper: Berkeley PDF)
- Thaler RH, Sunstein CR. Nudge: Improving Decisions About Health, Wealth, and Happiness. Yale University Press; 2008.
- Benartzi S, Beshears J, Milkman KL, Sunstein CR, Thaler RH, Shankar M, Tucker-Ray W, Congdon WJ, Galing S. Should governments invest more in nudging? Psychol Sci. 2017;28(8):1041-1055. doi:10.1177/0956797617702501
- Hummel D, Maedche A. How effective is nudging? A quantitative review on the effect sizes and limits of empirical nudging studies. J Behav Exp Econ. 2019;80:47-58. doi:10.1016/j.socec.2019.03.005
- Jachimowicz JM, Duncan S, Weber EU, Johnson EJ. When and why defaults influence decisions: a meta-analysis of default effects. Behav Public Policy. 2019;3(2):159-186. doi:10.1017/bpp.2018.43
- OECD. Behavioural Insights and Public Policy: Lessons from Around the World. OECD Publishing; 2017. (Figure A1: 200+ behavioural-insight units worldwide)