The Honor System Should Not Work
Every office with a communal coffee pot has tried the same doomed arrangement. Put out tea and coffee, leave a box, trust people to pay. The arrangement survives because it is simple. When payment is anonymous and enforcement is zero, economic theory predicts contributions will drift toward zero, and most departments quietly absorb the shortfall as the cost of caffeine.
Melissa Bateson, Daniel Nettle, and Gilbert Roberts at Newcastle University wondered if that drift is fixed. For ten weeks in 2006 they ran a field experiment in their psychology department coffee room with 48 staff as unwitting participants, tracking only money in and milk out. Each week they swapped the banner above the honesty box. One week a pair of eyes stared directly at the observer, varied in gender and angle but always gazing forward. Next week the banner showed flowers, same room and prices, same chance to underpay.
A Photocopy Tripled Payments
The team tracked cash in the box and litres of milk consumed, their proxy for drinks taken. Divide one by the other and the pattern leaps out. Eye weeks averaged 41.7 pence per litre versus 15.1 pence for flower weeks, a ratio of 2.76 to one, F(1,7)=11.55, p=0.011, explaining 63.8 percent of variance after log transformation. In practical terms, imagine a shared fridge used by two dozen people daily. Under flowers payment per litre fell to less than half of eye weeks, all from a six-inch-wide image swapped Monday morning before anyone arrived, a manipulation so subtle participants later reported no awareness of any experiment.
Why would printed eyes matter when everyone knows no one is watching? Bateson argued humans possess dedicated neural circuitry for gaze detection that fires automatically, triggering a low-level feeling of being observed even when we consciously know the eyes are fake, a feeling that activates the reputation psychology that normally keeps us honest when someone might judge us. Payment remained anonymous in fact but not in feeling, suggesting honesty depends on reputational cues rather than stable internal values.
An original scenario clarifies scale without claiming population size. Difference between conditions was £0.266 per litre (£0.417 minus £0.151). At a hypothetical 10 litres weekly, that equals £2.66 per week or about £138 per year for one coffee room, labelled explicitly as a scenario not an estimate of UK workplaces. Even a quarter of that gap would exceed the cost of printing posters.
The Replications Did Not Behave
Then meta-analyses arrived with sobering recalibration that every behavioral scientist should read before ordering posters. In 2013 Nettle and colleagues, including Bateson herself, pooled seven dictator-game experiments on watching eyes with combined n=887 and found watching eyes did not reliably increase mean donations, but did increase the probability of giving something at all, pooled odds ratio about 1.39, a modest shift turning some non-givers into minimal givers without making generous people more generous. In 2017 Northover, Pedersen, Cohen, and Andrews conducted two meta-analyses: roughly 26 experiments and about 2,700 participants for mean amount given, d ≈0.03, essentially zero, and roughly 27 experiments and close to 20,000 participants for whether people gave anything at all, d ≈0.13, with no convincing overall effect that artificial surveillance cues increase generosity.
Direct replications fared worse and often found no difference at all between eyes and control images. A 2016 online dictator game with 355 participants found F(2,353)=1.02, p=0.36, no effect, and eyes did not even increase self-reported feeling of being observed.
The Strongest Counterargument
Stated at full strength, the critique is that the 2006 result is real but highly conditional, emerging from 48 colleagues who knew each other, where a norm to pay existed and was enforced unevenly, and where weekly alternation made any change salient. In larger anonymous settings like train stations, parks, or tax letters, the effect shrinks toward zero when aggregated across dozens of preregistered replications with larger samples. Even the original statistics were contested: Carbon and Hesslinger (2011) argued the analysis violated assumptions, and when re-analyzed with four methods they deemed appropriate, only one remained significant.
If you print eyes in every hallway expecting a tripling of cooperation, you will be disappointed by the meta-analytic average, not confirmed by the original anecdote that went viral. The mechanism may be less about automatic gaze detection overriding rational calculation and more about reminding people of a local norm in a place where that norm already exists, a reminder that works only when novel, when no real people are present, and when decision is whether to contribute at all rather than how much, which is far narrower than the popular press version and explains why pilots failed when scaled to cities despite early enthusiasm from policymakers who had read only the original 2006 study and not the 27 replications that followed.
What We Didn't Prove
We did not track individuals. No one knows whether same people paid more in eye weeks or whether different people paid at all, because money and milk were measured at group level. Sample was 48 psychology staff and postgraduates at one UK university, hardly representative of public behavior. Ten weeks is short, carryover plausible: high-paying eye week may have left change that inflated next flower week. Milk is imperfect proxy; sugar, tea bags, personal cartons not measured. Weekly aggregates inflate effect sizes compared with person-level data, so any standardized effect derived from weekly means should be treated as an upper bound rather than a precise person-level estimate.
No preregistration existed in 2006. Flower control may depress prosocial behavior by priming non-social frame, making eyes look better by comparison. Demand cannot be fully excluded, though participants reported no awareness. Generalizability to high-stakes decisions remains untested. Coffee is cheap. Rent is not.
The Bottom Line
A cheap image of watching eyes increased honest payments 2.76 times in one kitchen, but pooled evidence from 27 later experiments suggests average effect on how much people give is near zero, with modest 39 percent increase in odds of giving something. Eyes work best when people know each other, when clear norm exists, and when few competing cues present. In those niches a poster may pay for itself. Outside them expect much less.
What You Can Do
Test the cue in your own shared resource before assuming it scales. Print two versions of your payment reminder, one with eyes gazing directly at viewer and one with neutral image matched for size and color, alternate weekly for at least eight weeks, track contributions divided by usage, log-transform weekly values to handle skew, run simple general linear model with image type as factor.
To apply psychology without poster, make real observability salient. Move box where people queue, add transparent tally with initials, or have named person thank contributors monthly. These create real reputational incentives and survive scrutiny that artificial eyes do not.
For personal habit change, reverse the lens. Notice where you are nudged by cues of being watched, from app permissions showing security badge to donation pages with live donor feed. Ask whether generosity in those moments reflects values or transient exposure, decide deliberately how much to give when cue removed.