How Drug Court Check-In Systems Enforce Accountability

Hawaii's HOPE probation program replaced rare, severe punishment with small, immediate consequences for missed check-ins, and a randomized trial found it worked. This traces the mechanics, the evidence, and what a phone app can and can't borrow from it.

In this article5 sections

Drug court check-in systems work by making the timing of testing unpredictable but the consequence for missing it fast and small: participants call a hotline or check an app most weekday mornings to learn whether they’re required to report that day, and a missed call-in or a positive test triggers an immediate, modest sanction rather than a delayed, severe one. The research suggests that structure, more than the severity of the eventual punishment, is what actually changes behavior.

The clearest test of that idea didn’t come from a psychology lab. It came from a courtroom in Honolulu in 2004, where a circuit court judge named Steven Alm got tired of watching the same offenders cycle through his docket.

The problem Judge Alm was trying to solve

Alm had noticed something that will sound familiar to anyone who has ever set a New Year’s resolution and broken it in February without consequence: the probation system he was working within almost never punished violations right away. A probationer would test positive for drugs, or skip a mandated appointment, and nothing would happen — not because the system didn’t care, but because revoking someone’s probation required a motion, a hearing, court dates, and often weeks of process. So violations piled up, quietly, until a probation officer finally filed a motion to revoke, at which point a judge would impose a severe sanction all at once, sometimes a multi-year prison term for an accumulation of infractions that had each, individually, gone unaddressed.

Alm’s insight was that this arrangement had the incentive structure backwards. People weren’t responding to the size of the eventual punishment; they were responding to the fact that, day to day, nothing happened. He redesigned his own courtroom’s probation supervision around a different bet: that certainty and speed would change behavior more reliably than severity ever had. He named it HOPE — Hawaii’s Opportunity Probation with Enforcement — and started running it in 2004.

How does a drug court check-in system actually work?

HOPE’s mechanics are simple enough to describe in a single paragraph, and the details matter because they’re exactly what makes the model different from ordinary probation. Each participant is assigned a color or code. Every weekday morning, they call a recorded phone line, and the line announces which colors are required to report for drug testing that day — assigned essentially at random, so nobody can plan around it. If your color is called, you go in that day for an observed test. If you miss the call-in, miss the required test, or test positive, the response is not a warning and not a referral to a future hearing — it’s an arrest, typically within days, followed by a short jail stay, often just a few days, imposed quickly and predictably. Then supervision continues. There’s no accumulation of unaddressed violations and no single catastrophic reckoning; the sanction for a given failure lands close enough to the failure itself that the two are legible as cause and effect.

Picture what that actually feels like from the participant’s side rather than the policy side. It’s 7 a.m. and you dial a number you’ve now dialed hundreds of times. A recorded voice reads off a short list of colors. Yours isn’t on it, so you hang up, and the day is yours. Tomorrow you’ll dial again, and you still won’t know the answer until you hear it — which is precisely the point. The randomness serves a real purpose: it keeps the system from becoming something a person can quietly learn to game.

That last sentence is really the whole theory in miniature: swift, certain, and fair — proportionate rather than maximal — is a stronger lever on behavior than severe, rare, and delayed. Most people’s intuitions about deterrence run the opposite direction. We tend to assume that the way to stop someone from doing something is to make the punishment for it as frightening as possible. HOPE’s premise, and the growing body of evidence behind it, is that once a consequence is large enough to register as real, adding more severity buys you very little, while adding certainty and speed buys you a lot. That’s a genuinely counterintuitive claim, and it deserves more attention than it usually gets, because most consumer accountability products — money bet on a habit, a fine for a missed workout — are still built on the opposite assumption, betting that a bigger number will do more work than a faster one.

What the randomized trial found

Alm’s redesign would be an interesting anecdote regardless, but it became something more than that because it got tested properly. Angela Hawken, a public policy researcher then at Pepperdine and now at NYU, teamed up with Mark Kleiman, a longtime scholar of drug policy and criminal justice at UCLA and later NYU, to run a randomized controlled trial comparing HOPE participants against probationers on standard supervision in the same court system. Their results, published around 2009, were stark by the standards of criminal justice research, a field where interventions routinely produce null or marginal effects. HOPE participants were substantially less likely to miss scheduled appointments, substantially less likely to test positive for drugs, and substantially less likely to be arrested for a new crime during the follow-up period, compared to the control group on ordinary probation.

What makes the trial worth taking seriously is as much about the design as the size of the effect. Randomization means the two groups started out comparable; whatever differences showed up afterward are harder to explain away as “the kind of people who succeed under HOPE would have succeeded anyway.” Kleiman went on to write about the swift-certain-fair framework at length, and the field’s national standard-setting body — historically the National Association of Drug Court Professionals, which now operates as All Rise — has incorporated the same logic into its published best-practice standards for problem-solving courts more broadly: frequent, unpredictable testing paired with immediate, modest responses to violations, rather than infrequent testing paired with severe, delayed ones.

Did it hold up everywhere it was tried?

This is the place to be careful, because a single successful trial in one courtroom is not the same thing as a portable, universal fix, and a full account of this story includes the part where replication got complicated. When HOPE-style programs were rolled out in other states and evaluated at scale — a multi-site study funded federally and carried out by RTI International in the mid-2010s — the results were noticeably weaker and more mixed than Hawaii’s original numbers. Some sites showed real improvements; others showed little difference from standard supervision at all.

The debate over why hasn’t fully resolved. One reading, advanced by Hawken and others close to the original research, is an implementation story: sites that didn’t reproduce the swiftness — where “immediate” sanction actually meant days or weeks, where call-in compliance wasn’t enforced with the same rigor Alm’s courtroom had — didn’t get the same results, because they weren’t really running the same program. A different reading is more skeptical of the underlying theory itself, suggesting that Hawaii’s original results may have benefited from conditions — a judge personally invested in the model, a smaller and more tightly managed caseload — that don’t travel easily to a larger bureaucracy. I don’t think the evidence fully settles which reading is correct, and I’d trust anyone who tells you it does a little less. What does seem well supported is narrower and more modest: the timing structure of consequences matters, probably a great deal, even if turning that insight into a program that survives contact with a different courthouse, a different budget, and a different set of line staff is its own separate and much harder problem.

What a phone app can borrow, and what it structurally can’t

It would be a mistake to treat this as a straightforward blueprint for consumer software, and the mistake deserves stating plainly rather than glossed over. A drug court can put someone in jail. It has subpoena power, sworn officers, and the machinery of the state behind every call-in. A habit-tracking app has none of that, and pretending otherwise would be dishonest about what’s actually on offer. Nobody using a wake-up or workout app is under a legal obligation to comply, and no missed check-in will ever result in anyone’s liberty being restricted. That asymmetry is real and it’s large, and it’s the honest boundary of the analogy, not a footnote to skip past.

What survives the transplant is narrower than the whole model — just the timing principle, while the enforcement power behind it stays entirely on the court’s side of the line. A missed commitment that produces some small, certain, near-immediate consequence changes behavior differently than a missed commitment that produces nothing until the rare moment someone finally cracks down — the same logic that shows up in comparisons between financial and social stakes, where the type of consequence matters as much as its size. DontSnooze is, in a narrow and voluntary sense, an attempt to borrow that same swift-and-certain-over-severe-and-rare principle for a consumer context: miss a check-in and something small and immediate happens — a person you chose gets notified — rather than nothing happening until a habit has quietly eroded for months. It’s opt-in, it’s reversible, and the stakes are social rather than custodial, which is a different animal entirely from a probation hotline. The design logic it borrows, more than the mechanism itself, is the interesting part: consequences that arrive close to the failure and don’t require a negotiation in the moment tend to hold up better than ones a person can talk their way out of, which is also the reasoning behind enforcement that doesn’t wait for permission once a deadline has passed.

None of this means an app can motivate the way a court order can, and any pitch that implies otherwise should be treated with some suspicion. What the HOPE research supports is a much smaller and more useful claim: if you’re designing any system meant to change behavior through consequences — a courtroom, a workplace policy, an app, a bet with a friend — the timing of the consequence deserves at least as much attention as its size. Judge Alm’s courtroom experiment, and the trial Hawken and Kleiman ran on top of it, is one of the better-documented pieces of evidence that the timing question isn’t a minor design detail. It may be closer to the whole game.

A caveat, because the analogy shouldn’t be oversold in the other direction either: an app that sends a photo to a friend and a probation hotline that sends someone to jail are not on the same continuum of severity, and treating them as differently-sized versions of the same thing would flatten a distinction that matters. What they share is narrower than that: a design choice about when a consequence lands relative to the failure. The size and nature of the consequence itself is where the comparison ends. That’s a smaller claim than “accountability apps work because of criminal justice research,” and it’s the only version of the claim this evidence actually supports.

Keep reading