Most of the Nudge Literature Disappears Once You Correct for Publication Bias
A 2022 meta-analysis found nudges work with a moderate effect (d = 0.43). A second team re-ran the same 200-plus studies through a model that accounts for publication bias and found the average effect can't be distinguished from zero, except in one domain. Here's which one, and why it matters for any app built on behavioral science.
In this article8 sections
In January 2022, one team of researchers published a meta-analysis in PNAS concluding that nudges work. Six months later, in the same journal, a second team ran the same pile of studies through a different statistical model and concluded that they mostly don’t. Neither team was wrong about their own math. They disagreed about how much to trust the underlying pile of studies in the first place, and the disagreement is a useful lens for reading any claim about behavior change, including the ones on this website.
The two papers, and where they actually clash
Stephanie Mertens and colleagues published a meta-analysis in PNAS covering more than 200 studies of nudges: changed defaults, reordered menus, framing tweaks, social-norm messages. Averaged across every study, the effect on behavior was moderate, d = 0.43, and the paper broke this down by domain: food-related nudges came in strongest at d = 0.72, financial nudges weakest at d = 0.25.
A team led by Maximilian Maier took the same body of studies and ran them through a different lens: a robust Bayesian meta-analysis, built specifically to account for one known distortion in published science. Studies that find an effect are simply more likely to get published than studies that find nothing, and reviewers, journals, and researchers themselves all contribute to that filter without anyone needing to act in bad faith. Maier’s model estimates what the average effect would look like if every unpublished null result that almost certainly exists were counted alongside the published ones. Run that correction across Mertens’ own 200-plus studies, and the overall effect can no longer be distinguished from zero.
Mertens’ team wrote a public reply. They didn’t dispute that publication bias exists in this literature, and conceded it plainly, but pushed back on treating “no distinguishable average effect” as the final word, pointing out that the studies vary enormously in what they’re measuring, which one flat correction can smear over. That pushback turns out to matter, because Maier’s own analysis found something the “nudging doesn’t work” headline usually leaves out: strong statistical evidence of publication bias in every domain Mertens measured except food, where the evidence for bias was much weaker. Food-related nudges are the one category the correction doesn’t clearly erase.
Why food and not finance
Neither paper fully explains the gap, but the pattern is suggestive rather than mysterious. Food-choice nudges (smaller default portion sizes, cafeteria layout, where the salad bar sits relative to the fries) tend to intervene at the exact moment of a low-stakes, repeated, visually immediate decision. Financial nudges (a preset retirement contribution rate, framing a bill as a loss versus a gain) usually target a decision made rarely, often abstractly, sometimes without the person even noticing a choice was made for them at all. A behavior repeated daily, with the intervention sitting directly in the visual path of the choice, appears to be a different kind of target than a once-a-year form nobody reads closely. That’s a hypothesis, not something either 2022 paper proves — nobody has run the head-to-head study that would confirm it.
Four questions, built from this one disagreement
Reading behavior-change claims for a living tends to produce one of two bad habits: taking every cited study at face value, or dismissing the entire field as unreplicable junk because one prominent debate went badly. Neither is calibrated. What the Mertens/Maier exchange actually supports is a narrower, more useful filter, four questions to ask before treating any single behavior-change statistic as settled:
- Has anyone modeled publication bias in this literature, not just cited a single study? A single RCT can be well-run and still sit inside a badly biased publication record. The question isn’t “was this study good,” it’s “has the whole shelf of studies like it been checked for a missing bottom half.”
- Does the claimed effect belong to a domain where corrected effects have actually survived, or one where they’ve evaporated? Food-adjacent, repeated, low-stakes daily choices are the one category with real post-correction evidence behind them so far. Financial, one-time, or abstract decisions are the category where the evidence disappeared.
- Is this a nudge (a preset choice with no real cost to ignoring it) or a commitment device (a pre-committed cost for failing)? These get discussed in the same breath constantly and they are not the same claim. A changed cafeteria layout and a $500 bet you lose for missing the gym work through different channels entirely, with different, separately studied literatures, and neither 2022 paper says anything about the second category at all.
- Would the person making the claim still make it if you asked them to bet money on a pre-registered replication? Not a rhetorical flourish: pre-registration and adversarial collaboration are the exact tools that produced this correction in the first place, and a claim’s author’s willingness to submit to that process is a reasonable, if informal, proxy for how much they actually believe their own number.
Applying this to accountability apps, including this one
DontSnooze, and every product like it, sits closer to the commitment-device side of question three than the nudge side. It isn’t a preset choice someone can quietly ignore; it’s a real cost, a friend getting notified, a photo going out, that only lands if the user fails to act. Neither 2022 paper’s finding, positive or negative, transfers cleanly to this category of product, which is a different kind of intervention than either one examined. The literature on commitment devices with real forfeitable stakes, going back to work by Dan Ariely and Klaus Wertenbroch on self-imposed deadlines, is older, smaller, and has its own separate replication problems that this piece hasn’t attempted to audit.
A real cost for failure doesn’t obviously reduce to the same “quiet option nobody notices” that corrected nudge effects seem to. That difference isn’t proof of anything on its own.
The older research on this exact category deserves naming directly rather than a passing gesture. Dan Ariely and Klaus Wertenbroch ran a set of experiments, published in Psychological Science in May 2002, asking three separate questions: would people voluntarily set themselves costly deadlines, would doing so actually improve their performance, and would they set those deadlines at the point that maximized performance. The answers were yes, yes, and no. People do impose real costs on themselves to fight procrastination, and it measurably helps. But left to choose their own deadline spacing, most people space it worse than an external deadline-setter would, which is itself a useful, humbling data point for any product that lets a user pick their own consequence schedule: the impulse to self-bind is real and helps, and people are still not great at building their own version of it.
The right level of confidence, applying the filter above to this blog’s own core claim, is closer to “plausible, and backed by an older and narrower body of work than the one that just took a hit” than “proven.” A reader who wants more than that should ask for it before taking any single accountability-app claim, including the ones made two paragraphs from here, at face value.
Running the filter on two more claims
The point of a four-question checklist is that it should work on more than one example. Two quick applications:
“Putting fruit at eye level in a cafeteria increases healthy eating.” This is a food-domain claim, the one category Maier’s own correction left mostly intact. Question two answers itself here: it’s exactly the kind of repeated, visually immediate, low-stakes decision where the corrected literature still shows an effect worth believing, with normal scientific caution.
“Automatically enrolling employees in a 401(k), with an opt-out instead of an opt-in, dramatically raises retirement savings.” This is a financial-domain preset-choice claim, the category with the weakest corrected effect in Mertens’ own numbers before Maier’s model was even applied. That doesn’t mean auto-enrollment does nothing: individual well-run studies of auto-enrollment on its own have shown large, durable effects, and this is a case where question one matters most, since nobody has yet checked the auto-enrollment literature specifically for the same bias, separate from the broad financial-nudge category it usually gets lumped into. This piece doesn’t know the answer, and neither claim above should be treated as settled by a two-paragraph mention.
This is also a useful discipline for headlines that never mention “nudge” or “publication bias” at all, including research-backed habit tricks like implementation intentions, which get cited constantly in productivity writing without anyone checking whether that literature has been bias-corrected the way the nudge literature just was. A news story announcing that some new app, supplement, or morning routine “increased follow-through by 40%” is making an empirical claim that the same four questions apply to, whether or not the word “nudge” appears anywhere in it. Most readers never see the underlying paper, only the rounded number a press release chose to lead with, and the press release is itself a form of publication bias operating one step past the journal: companies issue releases about their wins, rarely about their nulls.
What a convincing accountability study would actually require
A study that would actually settle this looks like something specific, rather than a vague “more research needed.” A convincing test of whether video-proof, social-consequence accountability actually changes wake-up behavior would need, at minimum: a pre-registered plan filed before data collection starts, so the hypothesis can’t quietly shift to match whatever the data shows; a randomized control group of people who want the same outcome but don’t get the social-consequence mechanism, rather than a comparison to people who never downloaded any app at all; an effect size reported with a confidence interval instead of a headline percentage; and, if it’s ever pooled with similar products in a future meta-analysis, the same publication-bias check Maier ran on the nudge literature, since a handful of app-company-funded studies showing success and no visible failures would be exactly the pattern that produced an inflated average the first time around.
DontSnooze has not run that study. What exists instead is aggregate usage data from people who chose to use the product, which can show correlation between check-ins and streak length but can’t rule out that the kind of person who signs up for social-consequence accountability in the first place is already more likely to follow through on things generally. That limitation is real, and it quietly weakens a large share of consumer-app “our users improved” claims across the entire category. It’s a different problem from the widely repeated 95% accountability statistic, which turned out not to trace back to a real study at all: a real, accurately reported number can still come from a study that was never built to rule out the obvious alternative explanation, which is a subtler failure than a number that was simply invented.
Where else this filter applies
This piece picked one well-aired disagreement in one adjacent field and used it to build a filter. It hasn’t run that filter against every claim on this website, and doing so would probably flag more than one of them. The four questions above are a starting discipline; a finished audit would take considerably longer, and the most useful thing a skeptical reader can do with them is point them back at whatever they just read, including this piece.
The same filter applies to a much smaller, more recent dispute over a single school-attendance postcard: a 2017 trial and a 2025 six-district replication of it disagree about which specific version of the message worked, for reasons that turn out to matter more than either study’s headline number. Question one and question two above apply just as directly to a postcard as to a cafeteria layout.
FAQ
Do nudges actually work, according to the research? It depends which correction you trust. Mertens et al. (2022, PNAS) found an average effect of d = 0.43 across 200-plus studies. Maier et al. (2022, PNAS) reanalyzed the same data with a publication-bias model and found no distinguishable overall effect, except in the food domain.
Why would publication bias make a real effect disappear? Studies that find a nudge works are more likely to get published than studies that find it doesn’t, which skews a simple average upward. A bias-correction model estimates what the average would look like if the missing unpublished nulls were counted.
Does this mean commitment devices and accountability apps don’t work either? Not necessarily — they’re a different category. A nudge quietly changes the option someone gets if they do nothing; a commitment device involves a real, pre-committed cost for failing, and neither 2022 paper examined that literature at all.
What’s a practical way to read a behavior-change study without a statistics background? Ask whether anyone has modeled publication bias in that specific literature, and whether the domain resembles one where corrected effects have held up (repeated, low-stakes, food-adjacent choices) or one where they’ve evaporated (one-time financial decisions).