Ranking Every Commitment Device by What It Actually Costs You
Commitment devices differ along two axes — what you stand to lose (money, social standing, or time) and whether enforcement is automatic or left to you — and that grid predicts which ones actually work. Financial devices with automatic enforcement and social devices with automatic enforcement both outperform anything self-policed.
In this article9 sections
Two things determine whether a commitment device holds when it matters: what a person actually stands to lose if they fail, and whether losing it is something they control in the moment. Most rankings of commitment devices skip the second question entirely, which is why so many of the “best commitment device” lists read the same and recommend things that fold under real pressure.
Building a Two-Axis Map Instead of a List
Ranking commitment devices on a single line, from “weak” to “strong,” hides the reason any of them work or don’t. Two devices can put the same locked-in cost on the line and land in completely different places, because the money means nothing if the person holding it is also the one deciding whether to hand it back to themselves.
So instead of a list, here’s a grid — call it the Stakes-Enforcement Grid. One axis is what’s at risk: money, social standing, or time and effort. The other axis is who enforces the consequence: the person themselves (self-policed) or an outside party or system that acts without asking permission (automatic). Every commitment device anyone actually uses can be placed somewhere on this grid, and the placement predicts, with reasonable accuracy, whether it survives contact with a tired, unmotivated version of the person who set it up.
The logic behind the enforcement axis isn’t new — it’s Schelling’s. In his 1984 paper “Self-Command in Practice, in Policy, and in a Theory of Rational Choice,” Thomas Schelling argued that self-imposed rules only hold when they’re what he called bright-line rules: no exceptions, no judgment calls, nothing to negotiate. A rule that allows “unless it’s really not a good day” isn’t a rule, because the exact moment you need the rule is also the moment you’re most motivated to invoke the exception. Automatic enforcement is Schelling’s bright line built into the system instead of into personal discipline — the exception clause simply doesn’t exist because no one with the authority to grant it is in the room.
Quadrant One: Money, Self-Policed
This is the weakest cell on the grid, and it’s also the most common piece of advice given to people trying to build a habit: “put money on it.” Venmo a friend $50 and tell them to keep it if you skip the gym. Set aside cash in an envelope labeled “lost if I sleep in.”
The problem is structural rather than about willpower. If the same person who’s supposed to lose the money is also the one who decides whether they lost — did I really skip the gym, or was that a rest day I’d already planned — the fine has no teeth. Reviews of how these devices are commonly built keep surfacing the same failure: a private financial pledge with no outside referee just becomes a thing you feel briefly bad about before deciding it doesn’t count this time.
Quadrant Two: Money, Automatically Enforced
Move the same stake to automatic enforcement and the picture changes substantially. StickK.com, cofounded by Yale economist Ian Ayres, is built almost entirely around this cell: users set a goal, put money behind it, name a referee, and — critically — the money moves without requiring the user’s continued cooperation. If a goal isn’t verified as met, the funds go to a charity, or worse, to an “anti-charity” the user finds actively distasteful. Ayres makes the case at length in Carrots and Sticks (Bantam, 2010): stakes only change behavior when there’s a real enforcement mechanism behind them, not just a stated intention.
A commitment-contract platform sitting in this quadrant only fails when the automation has a gap — a referee who rubber-stamps everything, or a verification step lenient enough to functionally return to Quadrant One. Side-by-side comparisons of these financial platforms against social ones tend to find that the financial ones work well right up until the enforcement gets soft, at which point they degrade to roughly the same failure rate as an unenforced pledge.
Money in this quadrant isn’t automatically safe, either. A well-known 2000 study of Israeli day-care centers found that adding a fine for late pickups made parents later, not earlier — the full mechanics of that specific backfire are worth reading on their own, but the short version is that pricing an obligation can strip out the guilt that was doing the actual enforcement work, which is a risk specific to layering a financial device on top of what started as a social one.
Quadrant Three: Social Standing, Self-Policed
This is the classic “tell everyone your goal” advice, and it’s a step up from a private financial pledge only because embarrassment is harder to quietly negotiate away than money is. But it’s still self-policed if the person under pressure is also the one who decides whether to report their own failure. Posting a goal publicly and then just… not posting an update when it goes badly is the standard failure mode, and it’s nearly frictionless, because nobody is checking.
Where this cell gets stronger is when the reporting itself becomes harder to dodge — a recurring group check-in with a habit of actually noticing absences, for instance. Research on how many observers a commitment needs before the social pressure becomes reliable suggests the number of people watching matters less than most people assume; what matters more is whether the group has a structure that catches silence, rather than depending on the person to volunteer their own bad news.
Quadrant Four: Social Standing, Automatically Enforced
This is the quadrant where the two hardest problems — real stakes, and stakes you can’t quietly opt out of — get solved simultaneously, and it’s the one most self-help advice skips because it’s harder to build than “tell a friend.” The consequence has to be social (something a person cares about because other people will see it) and it has to fire without the person’s cooperation, the same way an automatic bank debit doesn’t ask permission.
This is also the quadrant DontSnooze is built to occupy. Instead of a self-report — did you wake up, did you actually do the thing — failure triggers an automatic, unflattering notification to the person’s own friend group, with no step where the person decides whether to send it. That removes the exact failure mode from Quadrant Three: there’s no moment to quietly skip the confession. It’s fair to name the real limitation of sitting in this quadrant, though: an automatic social consequence is only as strong as the audience receiving it. If the friend group stops paying attention — a group chat that’s gone quiet, friends who’ve muted notifications — the automation still fires, but the social cost it’s supposed to deliver arrives with nobody there to register it. That’s a real weakness, and it’s specific to this quadrant rather than a knock against automatic enforcement generally; a financial device in Quadrant Two doesn’t have an equivalent “audience went quiet” failure mode, because a charity doesn’t need to be paying attention to receive a donation. Even accounting for that, a socially-staked device with automatic enforcement is a strong default for most people, because the two things it’s missing from Quadrant One and Three — a real cost, and no discretion to avoid paying it — are exactly the two things that predict whether any commitment device survives an actual bad morning.
Quadrants Five and Six: Time and Effort
The third axis of stakes — time or effort rather than money or reputation — tends to sit closer to the weak end regardless of enforcement, and it’s worth explaining why rather than just noting it. A self-policed time cost (say, “if I skip my workout I have to do double tomorrow”) suffers the same referee problem as a self-policed financial pledge. But even the automatically enforced version is weaker than its money or social equivalents, because time is fungible and recoverable in a way that spent money and damaged standing aren’t. Katy Milkman, Julia Minson, and John Beshears’s research on temptation bundling (Management Science, 2014) is instructive here by contrast: their intervention paired something people wanted to resist doing (exercising) with something they wanted to do anyway (listening to an addictive audiobook, available only at the gym). That’s not a stake at all — no cost is threatened — and it still moved behavior, which suggests time-based penalties are competing with a much easier, lower-friction alternative: just making the desired behavior itself more appealing, rather than punishing its absence.
Devices Can Migrate Across the Grid
The grid isn’t a fixed label glued to a product — the same device can migrate from one cell to another as its rules change, and watching that migration happen is often more instructive than picking a quadrant and stopping. Beeminder is the clearest example. A goal tracked with self-reported numbers and no penalty sits close to Quadrant One: money, self-policed, easy to fudge. Add a pledge that charges automatically on a missed data point, and the same goal slides toward Quadrant Two — money, automatic — without the user changing what they’re tracking at all. The one feature that does the most work in that slide is the one Beeminder built specifically to stop people from sliding back: an escalating penalty, so a habit of quietly raising your own goal right before missing it stops being free. Left flat, a financial penalty gets priced in and stops functioning as a deterrent the same way a familiar parking fine does; escalating it is a direct patch for the exact failure mode Quadrant Two is prone to.
The reverse migration happens too, usually by accident. A social pledge that starts in Quadrant Four — a friend group that’s notified automatically — can drift back toward Quadrant Three if the notification stops meaning anything, because everyone in the group has learned to scroll past it without comment. Automatic enforcement guarantees the message gets sent; it doesn’t guarantee anyone on the receiving end still treats it as a real consequence. That’s a genuinely different failure than the “audience went quiet” problem named above — it’s not that nobody’s there, it’s that being there has stopped costing the group anything either, and a witness who pays no price for witnessing eventually behaves like one who isn’t watching at all.
What the Grid Predicts, in One Sentence
A commitment device works in rough proportion to how real its cost is and how little discretion the person under pressure has over whether that cost lands — and the second variable, not the first, is the one nearly everyone underweights when they’re setting one up for themselves.
Choosing One for an Actual Habit
The practical use of this grid isn’t picking a permanent favorite quadrant; it’s diagnosing why a device someone already tried didn’t work. A private bet that fizzled almost certainly failed on the enforcement axis, not the stakes axis — raising the dollar amount won’t fix a device where the person holding the money is also the referee. Someone who’s tried public accountability and had it quietly lapse should look for a version where the reporting step is automated rather than self-initiated. The type of stake matters, but it’s the second question — who decides whether the cost actually gets paid — that separates the commitment devices people still remember keeping from the ones that faded out somewhere in February. Six of the field experiments this ranking draws on are worth reading directly, if only to see how thin the evidence gets for a couple of the devices ranked above.