Is It the Tracking or the Partner That Makes Accountability Work

Self-monitoring alone predicts success less reliably than monitoring paired with feedback, per Michie et al. (2009) and Leahey's 2018 study on what that feedback should sound like.

In this article7 sections

Self-monitoring alone — tracking a behavior with no partner, feedback, or review attached to it — is a markedly weaker predictor of reaching a goal than self-monitoring combined with at least one other self-regulation technique. That comes from a 2009 meta-regression of 122 published health-behavior intervention studies. A 2018 workplace weight-loss study points to one concrete form that “something else” can take in practice: a human accountability partner, and specifically the tone of what that partner says. Partners whose messages combined warmth with real pushback produced significantly larger results than partners who were only encouraging or only demanding, and both outperformed having no partner at all.

That framing cuts against a common shorthand version of the accountability story, where “tracking” and “having a partner” get treated as two separate, competing explanations for why accountability works. Two other pieces on this site cover whether accountability works at all, drawing on Gail Matthews’s 2015 goal-writing study and the Diabetes Prevention Program, and readers wanting the fuller research picture on why accountability works in the first place should start there. This piece asks a narrower question. Once tracking is already happening, what improves it: the mere presence of a partner, or the substance of what they say back?

What the largest study on self-monitoring found

Susan Michie, Charles Abraham, and three co-authors published a meta-regression in Health Psychology in 2009 titled “Effective Techniques in Healthy Eating and Physical Activity Interventions.” They analyzed 122 published intervention studies, coding each one for which specific behavior-change techniques it used, including self-monitoring, feedback on performance, and review of behavioral goals, and tested statistically which techniques predicted how effective the intervention was overall.

Self-monitoring came out as one of the strongest individual predictors in the entire set. That’s the part people tend to remember. The more specific finding is easier to miss: self-monitoring by itself was a considerably weaker predictor than self-monitoring combined with at least one additional technique from that list. Interventions that had people track their behavior but paired that tracking with nothing else did meaningfully worse than interventions that paired tracking with feedback or a structured review of the person’s goals.

Put plainly, tracking a behavior on its own is not where most of the studied benefit sits. It behaves more like a necessary ingredient that needs something else added to reach its full effect. Michie’s data doesn’t say what that “something else” has to look like between two real people; a meta-regression coding published papers for technique categories can’t get that granular. That’s where Leahey’s study picks up.

So what does the partner add?

Tricia Leahey, a researcher at the University of Connecticut, ran a study published in the Journal of Health Communication in 2018 called “The Buddy Benefit: Increasing the Effectiveness of an Employee-Targeted Weight-Loss Program.” The setup involved 704 employees enrolled in a 15-week online workplace weight-loss program. Roughly 54% of them opted into being paired with a buddy, another participant they’d check in with over the course of the program, while the rest went through the same program solo.

Buddy-paired participants lost significantly more weight and more waist circumference than solo participants. That result alone would support the simple version of the accountability story: having someone helps. But Leahey went further and coded the content of what buddies said to each other, sorting messages along two dimensions: acceptance (support, warmth, encouragement) and challenge (pushing back, holding the other person to their stated goal). The finding that matters most here is that buddies who scored high on both dimensions produced the largest reductions in BMI and waist size. Buddies who were only warm, or only demanding, did meaningfully worse than the combination, sometimes not much better than having no buddy at all.

Read against Michie’s framework, Leahey’s buddies map onto a specific slot. A partner combining warmth and pushback is a live, human version of feedback on performance, one of the technique categories Michie’s meta-regression flagged as strengthening self-monitoring’s effect. Leahey’s data adds a layer of resolution Michie’s broader synthesis couldn’t reach on its own: not all feedback works equally well. Feedback that’s only warm reads as permission. Feedback that’s only critical gets tuned out or resented. The combination is what moved weight and waist measurements in her data, which suggests the bar for that “something else” is narrower and harder to clear than simply having someone check in.

The two studies don’t line up as cleanly as they sound

It would be convenient to treat Michie and Leahey as two data points on the same line: monitoring plus something else beats monitoring alone, and here’s exactly what the something else should look like. That convenience should be resisted. These are not directly comparable studies, and treating them as though they confirm each other with precision would overstate what either one shows.

Michie’s meta-regression is a broad synthesis across 122 published studies on diet and physical activity, most of them not accountability-partner interventions at all; many used printed feedback reports, coaching calls, or automated reminders rather than an ongoing human relationship. The technique categories were coded after the fact from published descriptions, not measured directly as they happened. Leahey’s study, by contrast, is a single 15-week program in one population, corporate employees enrolled in a workplace wellness initiative, with buddies’ actual message content hand-coded along two specific interpersonal dimensions, measuring one outcome: weight and waist circumference. Calling a partner’s warm-and-challenging text message an instance of Michie’s feedback-on-performance category is a reasonable interpretive bridge, not something either paper tested directly. The two studies are compatible in spirit. They used different populations, different methods, and different outcome measures, and they don’t triangulate with the precision it would be convenient to claim.

Where the value likely comes from

Reading the two studies together, cautiously, and with that bridge acknowledged, suggests self-monitoring functions less like a complete solution and more like a necessary condition. Michie’s data says tracking on its own underperforms; something else has to be layered on top of it to reach its documented potential. Leahey’s data says that once a partner is the something else, the size of the benefit depends almost entirely on whether their messages do two jobs at once: staying warm enough that a person keeps showing up, while still being willing to say, in effect, that’s not what you told me you’d do.

A rough analogy: a smoke detector doesn’t put out a fire. It signals one. Whether that signal turns into a fire actually being contained depends on what happens next, a sprinkler firing automatically, or a person hearing the alarm and acting on it fast. The detector matters, and skipping it is far worse than having it, but detection by itself isn’t extinguishing. Michie’s finding is a version of this for self-regulation techniques: self-monitoring alerts a person to their own progress, but turning that alert into meaningfully better outcomes required, in her data, at least one additional technique layered on. Leahey’s buddies were one specific, human instance of that second layer, and the ones who combined warmth with real pushback did more of the extinguishing than the ones who only sounded the alarm.

This also bears on group size. Leahey’s design was strictly one-to-one, a single buddy rather than a group, and other research covered in this site’s piece on the Ringelmann ceiling in accountability groups suggests that once headcount grows past a small handful, individual signals get diluted and nobody feels specifically responsible for responding well. A message that would land as targeted challenge from one named buddy tends to dissolve into background noise in a group chat of twelve. If Leahey’s finding generalizes, it likely generalizes best in pairs rather than crowds, though that extension is inference, not something either study tested directly. Readers interested in the weight-loss literature beyond this single study can find a broader roundup in this site’s piece on accountability and weight loss research.

A limitation worth naming directly

None of this settles whether the combination Leahey found, high acceptance plus high challenge, generalizes outside weight loss to goals like waking up on time, finishing a degree, or quitting a habit. Weight loss has features that don’t map cleanly onto every goal: a scale gives a partner an unambiguous number to react to, in a way that “did you write today” or “did you get up on time” doesn’t always provide. It’s plausible that the acceptance-plus-challenge finding is specific to goals with clean, frequent, quantifiable feedback, and less applicable to goals that are more binary or harder to verify. That’s an open question, not a settled one, and it’s honest to say so.

What this means for an app like DontSnooze

DontSnooze works by having you send proof, a photo or a check-in, to a person you’ve chosen, on a schedule, with a real consequence if you don’t. It’s worth being honest about what that setup is and isn’t. Sending a photo to a specific person is self-monitoring plus a visibility layer, which is closer to a thin version of the “combine self-monitoring with something else” recipe Michie’s data points to than to Leahey’s full buddy dynamic, where two people exchange ongoing messages that mix encouragement with pushback over fifteen weeks. The app creates reporting and mild social stakes, but it doesn’t, by itself, manufacture the specific back-and-forth Leahey’s highest-performing buddies produced. If someone is hoping the app alone will replicate that dynamic without the two people involved engaging with each other’s replies, they’ll get less than the full effect Leahey documented. That’s a real gap, worth saying plainly rather than papering over. Even so, the recommendation holds, because Michie’s finding is the one that matters most here: self-monitoring alone underperforms, and the app’s core mechanic pairs monitoring with a real recipient rather than leaving it private, which is exactly the kind of pairing her meta-regression flags as doing more than tracking by itself. It’s well aimed at getting a person past the “alone” condition; what two people choose to say to each other on top of it is a second, separate opportunity that’s still theirs to take.

What a person can do with this

If this picture holds, the practical implication is straightforward, if a little deflating for solo-tracking habits: start tracking, because it’s the necessary condition none of this works without. But don’t expect tracking by itself to do the full job Michie’s data associates with the combined approach; look for, or build, a second piece on top of it, whether that’s a partner, a coach, or a structured review of your own goals. If a partner is available, the return on that partnership isn’t proportional to how often they check in; it’s proportional to whether their check-ins do both jobs at once. A partner who only says “you’ve got this” every morning and a partner who only says “you missed yesterday” are both leaving most of the available benefit on the table. The one who does both, in the same message, is rarer and harder to be, which may be exactly why it’s also the one the data says works best.

Keep reading