Oxford psychologists ran thousands of defensible analyses across three massive teen datasets instead of cherry-picking one, and the link between digital technology and adolescent wellbeing came out negative, real, and tiny: smaller than the link for wearing glasses, on par with eating potatoes.
Ask parents at the school gate and the story arrives pre-hardened: sometime around 2012, teenagers got smartphones, and teenage mental health fell apart. Haidt turned it into a bestselling book, the Surgeon General issued a 2023 advisory, and schools from New York to London began locking phones in pouches. The mechanism felt obvious, because every parent has watched it: the glowing rectangle, the vanished conversation, the kid who seemed fine until the phone arrived.
Amy Orben and Andrew Przybylski, psychologists at the University of Oxford, examined the same underlying data and found the obvious mechanism mostly missing; their 2019 Nature Human Behaviour paper is the largest and most methodologically paranoid test of the screen-time hypothesis ever run, covering 355,358 adolescents across three enormous datasets, two American and one British, each analyzed thousands of ways, because the authors refused to pick the analysis they liked best.
The method is the story: most papers pick one definition of screen time, one of wellbeing, one set of controls, and report whatever that single roll of the dice produces. Orben and Przybylski counted the defensible combinations instead, found more than 600 million, ran a large sample of them, and plotted every result using specification curve analysis, a technique whose whole purpose is to make cherry-picking impossible. Their curve exposed something else: slice the data however you like, and the screen-wellbeing link stays negative and very small.
At most, technology use explained 0.4% of the variation in teenagers' wellbeing, which the authors translated into same-dataset currencies. Regularly eating potatoes: nearly as negative, while wearing glasses ran one and a half times more negative; smoking marijuana 2.7 times more negative, and being bullied 4.3 times. On the positive side, sleep and breakfast carried far larger associations than technology use, consistent with the authors' earlier finding that they ran roughly three times as strong in the opposite direction. Orben's summary was blunt: tell me a teenager's screen time, and I still cannot predict their wellbeing.
What does 0.4% mean in human terms? The paper never runs this calculation; its numbers permit it. A variance share of 0.004 implies a correlation near 0.06, since 0.06 squared is 0.0036. Assume the two traits form a roughly bell-shaped cloud, pick two teenagers at random, and bet the heavier screen user reports lower wellbeing. Your win rate: one-half plus arcsin(0.06) divided by pi, about 0.52. A coin flip wins half the time. Screen time moves you two points past chance. The bell-shaped assumption flatters the screen-time hypothesis, since real survey data is messier, yet the conclusion survives: the variable the whole debate orbits is, for prediction, nearly inert.
Before it, a preregistered 2017 study of 120,115 English adolescents by Przybylski and Netta Weinstein found a Goldilocks curve: wellbeing peaked near one to two hours a day depending on the activity, then declined gently. Moderate engagement, they wrote, is not intrinsically harmful and may be advantageous; the 2019 paper asked whether that survived every defensible analytic choice, and it did.
The best case against this paper never disputes its arithmetic; it disputes its relevance. Teen mental health truly cratered around 2012, just as smartphones reached teenage pockets. CDC surveys show persistent sadness and suicidal thinking among American teens climbing steeply through the 2010s, while British self-harm admissions for young women rose in parallel; these are not fragile model artifacts but large, cross-national trends visible in raw data. So if screens explain four-tenths of one percent, the damage must flow through channels individual correlations cannot catch: whole peer groups losing sleep together, algorithmic feeds that arrived after most data was collected, network effects that rewire the social world even for light users.
The technical objection has teeth, and it appeared in the same journal. In 2020, Twenge, Haidt, Thomas Joiner, and Keith Campbell published six responses to the specification curve in one Matters Arising piece, arguing that individually defensible analytic decisions combined to shrink the estimates. Their exhibit: Kelly and colleagues' analysis of the same British dataset, where heavy social media users showed clinically relevant depressive symptoms at roughly twice the rate of non-users, which the critics called large enough to justify policy concern, adding that for girls one could conclude social media matters more than exercise or even heroin use. Orben and Przybylski replied that variance explained is the honest metric, because mean differences assume away third factors driving both heavy use and low mood; both sides published, the dispute stands unresolved, and the honest disagreement spans tiny to small, while no defensible analysis makes screens the dominant force the public debate assumes.
The authors' own 2022 follow-up concedes part of the point: Orben, Przybylski, Sarah-Jayne Blakemore, and Rogier Kievit followed 17,409 young people longitudinally in Nature Communications and found windows of sensitivity, with heavier social media use predicting lower life satisfaction a year later for girls aged 11 to 13 and boys 14 to 15, around puberty, and for both sexes at 19. The effects stayed small, and low satisfaction also predicted heavier later use at every age, undercutting any simple causal story. Still, the timing matters: one average across all of adolescence can dilute a real effect concentrated in a vulnerable window, and the data indicates early puberty, not the late teens.
The honest boundaries: this is correlational and cannot show screens cause anything, either way. Screen time was self-reported, and teenagers misreport like everyone else. Much of the data predates the algorithmic-feed era, reaching back in some cases to 2007. The study counts hours, not content: an hour video-calling grandparents equals an hour doomscrolling, which is like studying nutrition by weighing food. It cannot capture population-wide displacement, where every device alters the social environment for all, and it cannot rule out serious harm to vulnerable subgroups. That is precisely where the 2022 follow-up says to look.
Two things hold at once, and the panic demands you pick one: teen mental health deteriorated badly in the 2010s, and the smartphone arrived at the scene. Yet across 355,358 teenagers analyzed thousands of defensible ways, the individual-level link between technology and wellbeing is real, negative, and minute: smaller than glasses, on par with potatoes, dwarfed by bullying, sleep, and breakfast. The defensible reading is not that screens are harmless; it is that the hours variable, the thing parents monitor and schools confiscate, was never the main lever.
Stop optimizing the wrong number, because the paper's comparisons are a priority list in disguise. Sleep and breakfast carried the larger associations, so a phone rule that protects sleep is evidence-adjacent while a phone rule as moral theater is not. Ask what, not how long: the study counted hours, so its silence on content is total, and clinicians report that content and context outweigh the clock. Time your attention to the windows: early puberty, about 11 to 13 for girls and 14 to 15 for boys, is the peak-sensitivity window the 2022 data identifies, so presence matters most there. And if you make policy, take the authors at their word, because it is too small to warrant broad policy change; spend the capital on bullying, sleep, and the large, boring, unfashionable drivers of adolescent misery, not the 0.4% one.