Six Commitment-Device Experiments, and One That Turned Out to Be Fake

Real field experiments on precommitment, from Kenyan farmers to data-entry workers in India — plus the famous 2002 deadline study that Psychological Science retracted this month after evidence of data tampering.

In this article7 sections

Most explanations of how commitment devices work lean on the same two or three anecdotes for evidence. Here are six real field experiments that go further than the usual set — plus a seventh entry that belongs on this list for the opposite reason.

1. Farmers who wouldn’t buy fertilizer unless the offer expired that day

Esther Duflo, Michael Kremer, and Jonathan Robinson ran a field experiment with farmers in western Kenya (American Economic Review, 2011) offering a program called SAFI: a field officer visits right after harvest and offers a fertilizer voucher, at the normal price, with free delivery later in the season — but the farmer has to decide on the spot. Take-up was meaningfully higher than farmers simply being offered fertilizer at a discount later, when they had more time to think it over. The twist: farmers who used SAFI didn’t go on to buy fertilizer unprompted in later seasons, which told the researchers this wasn’t about farmers learning fertilizer was a good investment. It was about catching them at the one moment they had spare cash and hadn’t yet spent it on something else.

2. Audiobooks that only worked at the gym

Katherine Milkman, Julia Minson, and Kevin Volpp (Management Science, 2014) gave one group of participants iPods loaded with addictive audio novels — but locked to only play inside the gym. That group visited the gym 51% more often than a control group in the first weeks of the study, though the effect faded after Thanksgiving break. The number that stands out: by the study’s end, 61% of participants chose to pay to keep the restriction going voluntarily. People weren’t just tolerating having their own temptation weaponized against their procrastination. A majority wanted to keep paying for it.

3. The gym habit that only stuck for people who’d signed something

Heather Royer, Mark Stehr, and Justin Sydnor (American Economic Journal: Applied Economics, 2015) ran a four-week gym-attendance incentive at a Fortune-500 company. The incentive alone produced a strong short-term bump and a mostly forgettable long-term effect. The group that mattered was the subset offered a commitment contract on top of the incentive: their exercise gains were still detectable a full year after the financial incentive had ended. The lesson isn’t “incentives don’t work” — it’s that an incentive without a commitment mechanism attached seems to evaporate almost entirely once the money stops.

4. Workers who chose a pay structure designed to punish them

Supreet Kaur, Michael Kremer, and Sendhil Mullainathan (Journal of Political Economy, 2015) studied data-entry workers offered a choice of pay contracts, including “dominated” options that penalized low output on a given day but paid no bonus for high output — objectively worse than the alternatives on the table. Workers picked the self-punishing contract 36% of the time anyway, and it worked: using it boosted output by roughly as much as an 18% raise in the base piece-rate would have. People weren’t confused about the math. They were using a worse contract on purpose, as a tool against their own future slacking.

5. Bicycle-taxi drivers paid to stay sober during the day

Frank Schilbach’s field experiment with 229 cycle-rickshaw drivers in India (American Economic Review, 2019) offered small payments contingent on staying sober during working hours. Daytime drinking dropped substantially, with no change in how much drivers drank overall — the sobriety just moved to hours off the clock. The more interesting number: savings rose by roughly 50%, a bigger effect than income changes alone could explain, suggesting that staying sober during the day didn’t just protect wages, it protected the decisions the drivers made with those wages afterward.

6. The one that’s real, but doesn’t prove what it’s famous for

stickK’s own reported numbers — 78% goal completion for users with a financial stake and a referee, versus 35% for users with neither — get cited constantly as if they’re a controlled experiment. They’re not. They’re aggregated outcomes from a self-selected user base with no random assignment, published by the company whose product they describe. That doesn’t make the number worthless, but it’s a fundamentally different kind of evidence than the five peer-reviewed field experiments above it on this list, and the two shouldn’t get cited in the same breath as if they carry equal weight — a cost-by-cost ranking of the major commitment devices is a better place to weigh stickK’s own numbers against everything else on the market.

7. The famous one that wasn’t real at all

Dan Ariely and Klaus Wertenbroch’s 2002 paper on self-imposed deadlines — the study behind the widely repeated claim that people who set their own binding deadlines outperform people given open-ended ones — was cited well over a thousand times — estimates range from under 1,000 on Web of Science to over 2,000 on Google Scholar — and treated as a foundational result in this entire field. On September 2, 2026, Psychological Science retracted it. The retraction followed a failed replication and a forensic review by the research-fraud watchdog Data Colada, which reported it could not find an innocent explanation for the anomalies in one of the paper’s key datasets; Ariely has said his own records and memory can’t settle what happened. Whatever the full story turns out to be, the practical point for anyone citing “the deadline study” from memory is blunt: don’t, until something replaces it. A fair amount of commitment-device advice, on this site included, has absorbed a claim from a paper that no longer counts as valid evidence, which is worth correcting in public rather than quietly editing around.

Keep reading