Most Accountability Apps Are Optimized to Feel Effective, Not to Be Effective
Most accountability apps measure whether someone tapped a button or kept a streak alive, not whether they produced the outcome they were trying to get — and those two things drift apart over time. This piece uses research on deadlines and implementation intentions to explain why that gap exists and who it fools.
In this article7 sections
A streak, a completion percentage, and a “you did it!” notification all measure one thing: whether someone used the app. Whether they actually did the thing the app was supposed to produce is a separate question, and consumer accountability products routinely let the first stand in for the second.
Is there real evidence that public accountability changes behavior?
Some of it, and it’s more specific than the pitch most accountability apps make. In 2002, Dan Ariely and Klaus Wertenbroch published a study in Psychological Science using MIT students assigned to write three papers over a semester. Students who set their own deadlines did better than students given no deadlines at all — self-imposed structure beat no structure. But students given externally imposed, evenly spaced deadlines with real grade penalties for missing them did better than the students who set their own deadlines. The reason, according to Ariely and Wertenbroch, is that people are bad at spacing the deadlines they choose for themselves — left alone, they cluster their self-imposed checkpoints too close to the end, which defeats the purpose of having them.
Separately, the psychologist Peter Gollwitzer spent the 1990s studying what he called implementation intentions — specific if-then plans, like “if it’s 7am and the alarm goes off, then I get up and start the coffee,” rather than a general intention like “I’m going to wake up earlier.” His 1999 paper in American Psychologist summarized a body of experiments showing that people who formed these specific plans followed through at meaningfully higher rates than people who relied on intention or willpower alone, across tasks ranging from cancer screenings to resuming exercise after a break.
Read together, these two literatures point somewhere specific: what reliably improves follow-through is a real cost attached to a real deadline, or a plan specific enough to remove the decision point entirely. Neither paper is about announcing a goal to other people. Neither is about a number that goes up when you comply. They’re about structure — deadlines that bite, plans that pre-decide the moment — and that’s a narrower claim than “accountability works,” even though it gets marketed as the same thing.
Why does a streak counter behave like a game you can’t lose?
The claim to defend outright: a streak counter with no real stakes attached functions less like accountability and more like a scoreboard for a game where losing doesn’t cost anything. Not a weak form of accountability. A different category of object entirely, dressed in accountability’s clothing.
Think about what actually happens when a streak breaks. The number goes back to zero. Nobody else finds out unless you tell them. Nothing you own is at risk. No conversation gets harder tomorrow because of it. The consequence is entirely contained inside the app, and it consists of a digit changing — which is also, not coincidentally, the exact same event that happens when the streak continues, just in the other direction. A game where the penalty for losing is “the counter goes back to a number it’s been at before” is a game engineered to have players keep playing, not one engineered to produce an outcome outside itself.
Compare that to Ariely and Wertenbroch’s students, who lost actual points off a grade for missing a deadline, or to a Gollwitzer-style if-then plan, which doesn’t track anything at all — it just removes the moment where a decision could go the wrong way. Neither of those systems cares whether you feel good afterward. A streak counter cares about exactly one thing: whether you come back tomorrow and tap it again — a legitimate thing for a company to optimize for, and a separate thing entirely from accountability for the behavior underneath.
What’s the difference between tracking compliance and tracking the outcome?
This is where most of the argument actually lives, and it’s a measurement problem before it’s a psychology problem. The economist Charles Goodhart is credited with the observation, later sharpened by the anthropologist Marilyn Strathern into the version people actually quote: when a measure becomes a target, it stops being a good measure. Apply that to accountability apps and the pattern is everywhere. A streak measures whether you opened the app. A completion percentage measures how many check-ins you logged. A “you did it!” push notification fires the instant you tap a button, regardless of what preceded the tap. Every one of these is a proxy — a stand-in for the outcome you actually wanted, chosen because it’s easy for the app to observe. And once the proxy becomes the thing you’re managing, the underlying behavior can quietly detach from it. You can keep a fitness-app streak alive by logging a workout you cut short. You can keep a wake-up app’s completion rate perfect by tapping “done” from bed and going back to sleep. Neither action breaks the metric. Both defeat the point of tracking it.
That’s a related but distinct problem from the case that some accountability systems only look substantial while carrying no real weight, which is about whether anyone would notice, or anything would happen, if you quietly stopped. This is a step earlier than that: even when someone is watching and something does happen, the thing being watched can still be the wrong thing. An app can pass every test for “real” accountability — a friend gets notified, a cost is attached, someone would absolutely notice a gap — and still be measuring compliance with the app instead of the outcome the person signed up for.
Harder-to-fake evidence narrows this gap without closing it. A ranked look at which kinds of proof are easiest and hardest to fake makes the case that video beats a checkbox because video is expensive to falsify — which is true, and matters. But even a video only proves what happened in the ten seconds it captured. It answers “did this specific action occur at this specific moment,” which is a real question and a better one than “did you tap a button,” but it’s still not the same question as “did the day actually go the way you wanted it to.” Better proof shrinks the space between the proxy and the outcome. It doesn’t erase it, and treating “we require video” as a solved problem rather than a smaller version of the same one is its own kind of overclaiming.
Why do most apps ship the weak version anyway?
Because the weak version is dramatically easier to build and far less annoying to use, and both of those pressures point in the same commercial direction. A streak is a single number in a database that increments or resets — a solved problem, no third party required, no risk of upsetting a user badly enough that they delete the app. A real cost — money that actually leaves an account, a photo that actually reaches a friend, a deadline enforced by someone other than the person who set it — requires building an enforcement path, tolerating the complaints of the people it catches, and accepting that some fraction of users will churn specifically because the app did the thing it promised. Ariely and Wertenbroch’s own externally imposed deadlines only worked because the researchers were willing to actually dock grades. An app that flinches at the equivalent moment is running a different experiment than the one it’s borrowing credibility from, and reporting the results as if they transfer.
There’s also a retention story sitting underneath this, not as an accusation but as a plain fact about incentives: streaks, badges, and “you did it!” notifications are proven to bring users back tomorrow, independent of whether the underlying behavior happened today. That’s a real, useful product property on its own terms. It is also a completely different property from causing the outcome the user downloaded the app for, and a company optimizing for daily active users has every reason to build more of the first kind of feature and fewer of the second, even with the best intentions about the second.
Where else does this exact pattern show up?
Nowhere near accountability apps, usually, which is a reason to trust that it’s a real pattern and not something specific to habit-tracking software. A call center that grades agents on average call-handling time will get shorter calls — and a portion of those shorter calls will be agents hanging up before actually resolving anything, because the metric can’t tell the difference between a fast, competent answer and a fast non-answer. A school district that grades teachers on standardized test scores will get higher test scores, and some real fraction of that gain will come from time spent teaching the test’s specific format rather than the subject it’s meant to stand in for. In both cases nobody involved has to be lazy or dishonest for the gap to open up; they just have to respond, reasonably, to whichever number is actually being watched. An accountability app’s streak is the same shape of problem at a much smaller scale: the number is easy to observe, the underlying behavior is harder to observe, and effort quietly reallocates toward whichever one is actually being graded.
The reason to treat this as one pattern rather than three unrelated ones is that the fix is the same in every case, and it isn’t “try harder to measure honestly.” It’s closing the distance between the metric and the thing until faking the metric requires doing the real work anyway. A call center that grades resolution confirmed by a customer callback, a test that can’t be taught to directly, a wake-up check that requires producing something a still-asleep person can’t fake — all three are the same move, applied to different proxies.
Why doesn’t this work for everyone, even people who want accountability?
The argument that public accountability can backfire for shame-prone or highly autonomous people is about what watching does to a specific kind of person — it’s a claim about temperament and reactance. This is a different failure, and it doesn’t require a shame-prone user to show up: it happens to a confident, motivated person just as easily, because the number they’re checking every morning was never wired to the thing they actually wanted. Someone can be perfectly comfortable being watched, actually want the outcome, and still end up managing a proxy instead of a behavior, simply because the proxy is what the app put in front of them and the proxy is easier to keep green.
What would an app that measured the real thing actually look like?
Something closer to Gollwitzer’s if-then structure than to a streak: a specific trigger, a specific required action, and a cost attached to the action itself rather than to a record of it. A photo that goes to a friend group only when someone fails to produce something else in time works differently from a counter resetting, because missing the window has one specific, real consequence — a particular image landing in a particular inbox — instead of a quiet return to zero that nobody but the user will ever see.
A proof requirement closes the gap between “I claimed I did it” and “something happened at roughly the right time.” It doesn’t guarantee the deeper thing — that the morning was actually well spent, that the underlying goal is a good one, that the person even wanted this outcome versus feeling like they should. No accountability structure, weak or strong, answers that question. The strong ones are just honest about which narrower question they’re actually answering.
So the useful test isn’t “does this app hold me accountable.” It’s “what would actually be different today if I hadn’t done it.” If the honest answer is that a number would be one lower and nothing else in the world would change, that’s not a small accountability system. It’s a scoreboard for a game with no losing condition. Would knowing that change which one you check tomorrow morning?