π Social Science
Diversity Training Has an Evidence Problem. 418 Experiments, and 6 Tested the Real Thing
American companies spend roughly $8 billion a year on diversity training. A 2021 review of 418 prejudice-reduction experiments found that only six ever tested the workplace kind of training, and the two largest of those found an average effect near zero.
- Title
- Prejudice Reduction: Progress and Challenges
- Authors
- Paluck EL, Porat R, Clark CS, Green DP (2021)
- Institution
- Princeton University; Hebrew University of Jerusalem; Columbia University
- Journal
- Annual Review of Psychology, 2021, 72, 533–560
- Sample
- 418 experiments reported in 309 manuscripts, published 2007–2019 (experimental studies with random assignment only)
- Method
- Systematic review with meta-analysis of average effects by theoretical approach, plus qualitative review of landmark field studies
- Key Finding
- 76% of experiments tested light-touch interventions under ten minutes; the pooled effect (d = 0.357) fell 48% (to d = 0.187) among the largest studies; only six experiments tested actual workplace diversity training
- Effect Size
- d = 0.357 overall, dropping to d = 0.187 in the top quintile of sample sizes (85% overlap between treated and control groups at the pooled estimate). Diversity-training subset: d = 0.30 overall, d = 0.07 in the two largest studies. Unmistakable publication bias in every theoretical domain
- Counterintuition
- ⚡⚡⚡⚡ 4/5
- Replication
- Meta-analyzed (418 experiments). Conclusions converge with independent reviews (Forscher et al. 2019 on implicit-bias procedures; later DEI evidence reviews). No retraction, no correction, no PubPeer flags as of October 2026
When a company's bias makes headlines, the corporate response now runs on rails. A public apology, a task force, and mandatory training for everyone. American businesses spend roughly $8 billion a year on diversity training, according to a widely cited Harvard Business Review figure; the sessions have become as standard as fire drills, and the premise underneath all of it is rarely examined: that a workshop changes how people behave at work.
A team of researchers decided to examine it. Elizabeth Levy Paluck of Princeton, Roni Porat, Chelsey Clark, and Donald Green of Columbia spent years assembling every experimental test of prejudice reduction published between 2007 and 2019. Their review, published in the Annual Review of Psychology in 2021, covers 418 experiments reported in 309 manuscripts, restricted to studies that randomly assigned people to treatment and control. It is the largest quantitative assessment the field has ever produced. Its verdict is uncomfortable.
Start with what the literature mostly contains. Seventy-six percent of the experiments tested what the authors call light-touch interventions: brief exercises, under ten minutes, cheap to deliver, and assumed to have lasting effects. The modal study asked college students to imagine a positive encounter with an outgroup member, or to complete a short perspective-taking writing exercise, then measured their attitudes minutes later. Only 8 percent of the 418 experiments were preregistered or shared open data, which means the literature grew up largely without the safeguards that keep researchers honest about null results.
The headline number looks respectable at first. Pooling everything, the average intervention reduced prejudice by d = 0.357, a modest but meaningful shift. The authors translate it: after the intervention, there would be 85 percent overlap between treated people and controls. Then they apply the oldest quality check in meta-analysis: sort studies by sample size and look at the biggest ones, which are hardest to game and most precise, and watch what happens to the average effect once the small, noisy studies most vulnerable to file-drawer bias fall away. It drops 48 percent, to d = 0.187. On a feeling thermometer, that is five times smaller than the warming of American attitudes toward gay people over the past two decades, and in every theoretical domain the authors examined, without exception, smaller studies reported bigger effects. That pattern has a name, publication bias, and the authors call the evidence for it unmistakable: the studies that reach journals are disproportionately the ones with flattering results, while the disappointments stay in file drawers.
Now the part that matters for the $8 billion. Paluck's team counted only studies testing programs that called themselves diversity, sensitivity, or cultural-competence training, the kind companies actually buy. Across the entire decade, they found six experiments with a pooled effect of d = 0.30, but four of the six ran on university students, and those campus studies, averaging d = 0.45, did all the lifting. Restrict the sample to the two studies in the largest-sample quintile and the average falls to d = 0.07, essentially zero. Read that again.
One of those two is the study the authors single out as a "quantum leap". In 2019, Edward Chang and colleagues at Wharton ran a preregistered field experiment with 3,016 employees of a global professional-services firm, testing an hour-long online diversity training against an active placebo control and tracking real workplace behavior for 20 weeks. It remains the only experiment to test diversity training as it is naturally implemented, inside a real company. The results were mixed, exactly as the paper's title promises. Among employees whose attitudes toward women were relatively less supportive, the training changed attitudes but not behavior. Among women and minorities, it changed behavior: American women who took the training became more likely to invite other women to mentoring coffees, about one extra invitation for every five women trained. On the racial-bias module, employees became more willing to admit bias, with no measurable behavior change. The training moved the people who needed it least.
Here is the arithmetic the review leaves to the reader. Eight billion dollars a year in corporate spending, against six experiments in a decade on the actual product, which works out to roughly $1.3 billion in annual spending for every experimental study ever conducted on the workplace kind of diversity training. Science is not funded per study, and the comparison is deliberately crude, but it captures the asymmetry the authors document: an industry operating at national scale on an evidence base you could read in an afternoon.
The strongest case against this bleak picture starts with a rival meta-analysis. In 2016, Katerina Bezrukova and colleagues pooled four decades of diversity-training research in Psychological Bulletin and found positive effects on attitudes, and sometimes behavior. Paluck's team answers that the earlier review counted almost anything as diversity training, including, in one striking case, Rwandan radio soap operas about ethnic reconciliation. Both findings can be true at once: a broad tent of interventions shows promise, while the narrow tent of corporate workshops has almost no experimental backing. A second counterargument is stronger still. Intergroup contact genuinely works. Thomas Pettigrew and Linda Tropp's 2006 meta-analysis of 515 studies found that contact between groups typically reduces prejudice, with the most rigorous studies showing the largest effects, and field experiments since, from door-to-door canvassing to mixed soccer teams in Iraq, keep confirming it, which is why the review's landmark section returns to contact again and again as the rare approach with genuine field evidence behind it. The review itself spotlights these landmark studies, and its claim is not that nothing works, only that the thing companies buy in bulk, the brief workshop, is not where the evidence lives.
There is a third nuance worth stating plainly, because it cuts against the review's own framing. Changing minds and changing behavior are different projects, and a separate 2019 meta-analysis of 492 studies by Patrick Forscher and colleagues found that procedures shifting implicit-bias scores produced what the authors called trivial changes in behavior, with no evidence that attitude change mediated behavior change, which means that if the goal is fewer discriminatory decisions rather than warmer survey answers, the workshop model is aimed at the wrong target.
The review's conclusion is blunt even by academic standards. Much of the research effort, the authors write, is "theoretically and empirically ill-suited to provide actionable, evidence-based recommendations for reducing prejudice." That sentence has been widely cited, and later reviews have converged with it rather than rebutted it, while the field's response, to its credit, has been to demand preregistration, field partnerships, and behavioral outcomes measured months later, exactly the template Chang's study set.
What We Didn't Prove
- This is a review of experiments published from 2007 to 2019. It cannot speak to training programs designed after the window closed, or to effects the literature never measured.
- The d = 0.07 figure for large diversity-training studies rests on two studies, so the honest reading is not "the true effect is 0.07" but "we barely know."
- Nearly all outcomes were attitudes captured on surveys, often minutes after the intervention. Behavior inside real workplaces is the outcome that matters most and the one least studied.
- The literature is predominantly American and college-student heavy. These results may not travel to other countries or to non-student adults.
- Publication bias means true average effects are probably smaller than the published numbers, but the exact shrinkage cannot be recovered from published data alone.
- The $1.3-billion-per-study arithmetic is illustrative, not an economic analysis. Corporate spending also buys legal compliance, delivery, and signaling, not just attitude change.
- The review does not test whether mandatory training can backfire. Other researchers argue it can, but that claim rests on observational company data, not on these experiments.
The Bottom Line
A decade of experiments, 418 of them, shows that prejudice-reduction interventions mostly nudge survey answers a little, briefly, in small studies, while the corporate workshop at the center of an $8 billion industry has six experiments behind it, with the two largest averaging near zero. What does work, sustained contact between groups under cooperative conditions, looks nothing like a one-hour video. The evidence does not say training is useless. It says we are buying the wrong thing with almost no receipt.
What You Can Do
- If you run a company, stop grading training by smile sheets. Ask vendors for behavioral outcomes measured months later, or run a randomized pilot of your own, the way Chang's corporate partner did.
- Prefer structure over sessions. Across literatures, the consistent finding is that cooperative contact, mixed teams with shared goals, moves behavior more than any workshop.
- Make attendance voluntary where you can. Mandatory programs aimed at skeptical employees carry the weakest evidence and the strongest backlash risk.
- Measure what matters: hiring, promotion, retention, and pay gaps by group, tracked over years. Attitude surveys are the easiest metric to game.
- If you sit through the training, treat it as the start of a conversation rather than a cure. Durable change comes from repeated contact and accountability, not a single hour.
- If you fund research, the review's plea is explicit: preregister, partner with organizations, and measure real behavior. Only 8 percent of a decade of experiments did the first.
Sources
- Paluck EL, Porat R, Clark CS, Green DP. Prejudice Reduction: Progress and Challenges. Annu Rev Psychol. 2021;72:533-560. doi:10.1146/annurev-psych-071620-030619
- Chang EH, Milkman KL, Gromet DM, Rebele RW, Massey C, Duckworth AL, Grant AM. The mixed effects of online diversity training. Proc Natl Acad Sci USA. 2019;116(16):7778-7783. doi:10.1073/pnas.1816076116
- Bezrukova K, Spell CS, Perry JL, Jehn KA. A meta-analytical integration of over 40 years of research on diversity training evaluation. Psychol Bull. 2016;142(11):1227-1274. doi:10.1037/bul0000067
- Pettigrew TF, Tropp LR. A meta-analytic test of intergroup contact theory. J Pers Soc Psychol. 2006;90(5):751-783. doi:10.1037/0022-3514.90.5.751
- Forscher PS, Lai CK, Axt JR, Ebersole CR, Herman M, Devine PG, Nosek BA. A meta-analysis of procedures to change implicit measures. J Pers Soc Psychol. 2019;117(3):522-559. doi:10.1037/pspa0000160
- Dobbin F, Kalev A. Why Diversity Programs Fail. Harvard Business Review. July-August 2016. hbr.org/2016/07/why-diversity-programs-fail
- Broockman D, Kalla J. Durably reducing transphobia: a field experiment on door-to-door canvassing. Science. 2016;352(6282):220-224. doi:10.1126/science.aad9713