TL;DR: Accountability is not one thing, it is five separate functions, and most arguments about habit tracking fail because they compare whole products instead of functions. The best meta-analysis on the subject, 138 experiments covering 19,951 people, found that monitoring your progress does move goal attainment (d = 0.40) and that the effect gets larger under two specific conditions: when you physically record the result, and when you report it to another person. An AI assistant is extremely good at four of the five functions and structurally incapable of the fifth, because reporting to a system that cannot think less of you is not reporting to another person. The first randomized field study of an LLM habit coach, run over four weeks with 54 participants, found exactly the pattern this predicts: better beliefs, more enjoyment and more self-compassion in the AI condition, with no short-term advantage in actual activity. Meanwhile a workplace field experiment showed that adding a small financial commitment, taken up by only 12% of the eligible group, produced behavior change still visible a year later. The practical conclusion is not “AI coaching does not work.” It is that you should use the AI for the four functions it dominates and buy the fifth somewhere else.
There is a specific disappointment that most people who have tried an AI habit coach will recognize. The conversation is good. It asks better questions than a habit app ever did. It notices that you skipped Thursday and it does not shame you about it. You finish the exchange feeling genuinely more capable.
And then Friday goes exactly the way Thursday did.
This is not a failure of the model, and it is not a failure of your discipline. It is a structural gap that the evidence identified years before anyone built an LLM coach, and it is entirely fixable once you can see it clearly.
The mistake is comparing products instead of functions
Ask which is better, a paper journal or an app or an AI assistant, and you get an unanswerable question, because each of those is a bundle of unrelated capabilities sold as one thing. The useful move is to unbundle.
Strip habit tracking down and it performs five distinct jobs:
- Cue and timing. Deciding in advance when and where the behavior happens, and firing at that moment.
- Recording. Capturing whether it happened, in a form you can see later.
- Feedback and interpretation. Turning the record into a judgment about what is actually going on.
- Social stake. Someone whose opinion you care about will know the answer.
- Repair. Getting back on after a lapse, before the lapse becomes the new pattern.
Every system on the market does some of these well and quietly leaves others to you. Once you see which is which, the choice stops being a matter of taste.
What the evidence actually establishes
Before scoring anything, here is the underlying research, with the numbers as reported.
| Mechanism | Source | Design | Result as reported |
|---|---|---|---|
| Monitoring goal progress | Harkin, Webb, Chang, Prestwich, Conner, Kellar, Benn and Sheeran, Psychological Bulletin, 2016, 142(2) | Meta-analysis of 138 randomized studies, N = 19,951 | Interventions raised monitoring frequency (d = 1.98) and goal attainment (d = 0.40, 95% CI 0.32 to 0.48). Effects were larger when progress was physically recorded and when results were reported to another person. |
| Planning the cue | Gollwitzer and Sheeran, Advances in Experimental Social Psychology, 2006, volume 38 | Meta-analysis, 94 independent tests | Implementation intentions (“if situation Y, then I will do X”) produced a medium-to-large effect on goal attainment, d = 0.65. |
| Commitment with a stake | Royer, Stehr and Sydnor, American Economic Journal: Applied Economics, 2015, 7(3) | Workplace field experiment, Fortune 500 employees, gym attendance tracked for a year | Baseline about 20% attending weekly; the one-month incentive roughly doubled that. Incentive alone: only 25% of the gain persisted in month one and most was gone by month two. Incentive plus an optional self-funded commitment contract (12% take-up): about 30% attending weekly in months two and three against 20% for control, with effects still detectable a year after the incentive ended. |
| How long automaticity takes | Lally, van Jaarsveld, Potts and Wardle, European Journal of Social Psychology, 2010, 40(6) | 96 volunteers, one daily behavior, 12 weeks of self-report | Median 66 days to reach 95% of asymptotic automaticity, individual range 18 to 254 days. The curve was modelled for 62 participants and fitted well for 39 of them. |
| An LLM coach against a no-LLM control | Jörke, Genc, Teutschbein, Sapkota, Chung, Schmiedmayer, Campero, King, Brunskill and Landay, Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems | Four-week randomized field study, N = 54, physical activity app with and without an LLM coach | Both conditions significantly increased activity and doubled the proportion of participants meeting recommended weekly guidelines. The LLM condition reported stronger beliefs that activity was beneficial, greater enjoyment and more self-compassion, but showed no descriptive advantage in short-term activity levels. |
Read those five rows together and a shape emerges that no single study states outright.
Tracking works, but modestly. Planning the cue works better than tracking the outcome. A stake outperforms both, and it is the only intervention on the list whose effect was still measurable a year later. The timeline is long enough that the stake has to survive months, not weeks. And the newest entry, the LLM coach, moved the psychological layer without yet moving the behavioral one.
Two moderators do most of the work
The Harkin meta-analysis is the closest thing this field has to a settled result, and its two moderators are the whole story.
Physical recording beats a tap. Monitoring had larger effects when people physically recorded their progress. This is uncomfortable for the entire category of frictionless tracking, because frictionless is the product promise. A wearable that logs your steps without asking you anything has removed the exact step the evidence says carries weight. It has not made tracking easier, it has made tracking passive, and passive is a different intervention.
Reporting to another person beats reporting to yourself. This is the moderator that decides the AI question. Not “does the AI ask about it,” not “does the AI remember,” but: is there a person whose picture of you changes based on the answer?
An AI assistant fails this test by construction, and the failure is not a limitation that a better model closes. The mechanism requires a mind that will hold a revised opinion of you tomorrow. A model that begins each conversation without stakes in your reputation cannot supply that, no matter how well it simulates concern. Anyone who has used one knows this at some level, which is why telling a chatbot you skipped the gym costs nothing and telling a training partner costs something.
The five-function scorecard
Here is the framework applied. Scores are 0 (absent), 1 (partial) or 2 (strong), assigned by reading each system against the evidence above. This is a CEOtudent editorial framework, not a measured result: the underlying studies are real, the scoring is our judgment about which system delivers which function, and you should feel free to disagree with an individual cell.
| System | Cue and timing | Recording | Feedback | Social stake | Repair | Total (of 10) | Coverage |
|---|---|---|---|---|---|---|---|
| Paper log or bullet journal | 1 | 2 | 1 | 0 | 1 | 5 | 50% |
| Habit app with streaks | 2 | 1 | 1 | 0 | 0 | 4 | 40% |
| Wearable, passive tracking | 0 | 2 | 1 | 0 | 0 | 3 | 30% |
| Human partner or coach | 1 | 1 | 2 | 2 | 2 | 8 | 80% |
| AI assistant, 2026 | 2 | 1 | 2 | 0 | 2 | 7 | 70% |
| AI assistant plus a human stake | 2 | 2 | 2 | 2 | 2 | 10 | 100% |
The reasoning behind the less obvious cells:
- Paper scores 2 on recording because physical recording is the moderator Harkin identified, and paper is the only system on the list where recording is unavoidably physical.
- The streak app scores 0 on repair. A streak is a repair penalty disguised as a motivator: the moment you break it, the app’s central feature is now a monument to your failure. Given a median of 66 days to automaticity with a range running past 250, a mechanism that punishes the inevitable first miss is working against the timeline it claims to serve.
- The wearable scores 0 on cue because it does not ask you to decide when and where anything happens. It observes. Gollwitzer and Sheeran’s d = 0.65 belongs to the system that made you plan, and passive tracking never does.
- The AI scores 2 on repair, and this is its genuine and underrated strength. The self-compassion finding in the CHI 2026 study is not a soft consolation prize. Lapse recovery is where most habits actually die, and an interlocutor that responds to a missed Thursday with a recalibration rather than a guilt spiral is doing real mechanical work.
- The AI scores 0 on social stake, and this single zero explains the Bloom result. Better beliefs, better enjoyment, better self-compassion, no behavioral advantage. Four functions strong, the fifth missing.
The gap has a price and the price is known
The useful thing about the Royer field experiment is that it puts a number on what the fifth function is worth.
Incentives alone: 25% of the gain surviving one month, essentially nothing by month two. Incentives plus a real stake: attendance elevated through the commitment period and still detectable a year later. The take-up rate matters too, and it is humbling. Only 12% of the eligible group chose to create a commitment contract at all. Most people, offered the mechanism that works, decline it.
Which is the honest reason AI coaching is popular. It is the accountability that costs nothing, and the thing that costs nothing is the thing that does not bind.
What to actually build
Use the AI for the four functions where it beats every alternative, and buy the fifth deliberately.
1. Let the AI write the implementation intention, not the goal. The highest-effect intervention in the evidence base is the least glamorous. Do not ask for a plan to get fit. Ask it to convert one behavior into a strict if-then with a named time, a named place and a named trigger, then to interrogate the version you produce for vagueness. Anything that lacks a specific cue is not an implementation intention, it is a wish with a deadline.
2. Keep recording physical, or at least effortful. This is the one place to ignore the product design of the last decade. A written mark takes four seconds and it is the version the meta-analysis rewards. If you will not use paper, make the digital entry require a sentence rather than a tap.
3. Give the AI the interpretation job, because it is genuinely good at it. Feed it three weeks of raw records and ask what pattern it sees, which day fails most, what preceded each miss. This is the function where it clears every other system in the table, and it is the one people use least.
4. Buy the social stake from a human. One named person, one fixed weekly moment, one number reported. It does not need to be a coach and it does not need to be reciprocal. It needs to be a person who will notice.
5. Add a stake with teeth if the human is not enough. The Royer design is trivially reproducible: money at risk, forfeited to a cause you dislike, over a defined window. It works and almost nobody does it, and those two facts are related.
6. Set the review at 66 days, not 21. Automaticity in the Lally data had a median of 66 days and a range out to 254. Any system that evaluates you at week three is evaluating you before the mechanism has run.
For a deeper look at how long the underlying curve actually takes, see our piece on how long it takes to build a habit. If you are building the surrounding routine rather than a single behavior, what the evidence says about morning routines covers which components survive scrutiny, and energy auditing covers the input side. For the measurement question specifically, which wearable metrics actually mean anything applies the same evidentiary standard to the passive-tracking column of the table above.
The CEO and the student read this differently, and both are right
Run it as a CEO and the question is not motivation, it is governance. You would never accept a reporting line where the report goes to a system that has no opinion and no memory of your last commitment. You would insist on a named counterparty and a fixed cadence. Apply that standard to yourself and the fifth column stops being optional.
Run it as a student and the question is what you are learning. The AI’s interpretation function is a genuinely new capability: nobody in 2015 had a patient analyst willing to read three weeks of messy self-report and name the pattern. Use it for that. Just do not confuse having understood the pattern with having changed it.
FAQ
Will a better model close the social-stake gap?
Not by getting better at the conversation. The moderator in the evidence is reporting to another person, and the operative word is person, meaning someone who carries a revisable opinion of you into tomorrow. A more capable model produces a more convincing simulation of concern, which may make the gap less noticeable without making it smaller. The realistic path is not a better model but a different architecture: an AI that reports to a human on your behalf, at your instruction, is a genuinely different mechanism from an AI that talks to you.
Does that mean AI habit coaching is not worth using?
No, and the table says why. Seven of ten functions, including the two that most systems handle worst, cue design and lapse repair, is a strong result. The error is expecting the remaining three points to arrive from the same place.
Is the Bloom study strong enough to conclude this?
Not on its own, and it should not be treated that way. It is one four-week study with 54 participants, and the authors themselves frame the psychological shifts as possible precursors to longer-term change rather than as a null result for LLM coaching. Four weeks is also well short of the 66-day median in the Lally data, so a study of that length may be measuring the wrong window. What makes it worth citing is that its pattern matches what the much larger Harkin meta-analysis would predict, and converging evidence from independent designs is stronger than either piece alone.
What if I do not have anyone to report to?
Then the financial commitment is your substitute, and it is the better-evidenced of the two anyway. It also has a specific advantage: it does not require anyone else to remember, show up, or care, which is where informal accountability partnerships usually fail.
Is streak-based tracking actively harmful?
The evidence does not support that strong a claim, and we are not making it. What the evidence supports is narrower: streaks add nothing to the two moderators that matter, and their central mechanic imposes its largest psychological cost at the exact moment, the first miss, when the timeline data says a lapse is statistically normal. Use them if they help you, but do not count them as accountability.
Sources
- Harkin, Webb, Chang, Prestwich, Conner, Kellar, Benn and Sheeran, Does Monitoring Goal Progress Promote Goal Attainment? A Meta-Analysis of the Experimental Evidence, Psychological Bulletin, 2016, volume 142, issue 2
- Gollwitzer and Sheeran, Implementation Intentions and Goal Achievement: A Meta-Analysis of Effects and Processes, Advances in Experimental Social Psychology, 2006, volume 38, pages 69 to 119
- Royer, Stehr and Sydnor, Incentives, Commitments and Habit Formation in Exercise: Evidence from a Field Experiment with Workers at a Fortune-500 Company, American Economic Journal: Applied Economics, 2015, volume 7, issue 3, pages 51 to 84
- Lally, van Jaarsveld, Potts and Wardle, How Are Habits Formed: Modelling Habit Formation in the Real World, European Journal of Social Psychology, 2010, volume 40, issue 6, pages 998 to 1009
- Jörke, Genc, Teutschbein, Sapkota, Chung, Schmiedmayer, Campero, King, Brunskill and Landay, Bloom: Designing for LLM-Augmented Behavior Change Interactions, Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, Stanford University
- Wood and Neal, A New Look at Habits and the Habit-Goal Interface, Psychological Review, for the context-dependency account of habit that underlies the cue-and-timing function
This content was compiled with the support of AI following in-depth research, then written and prepared for publication by the CEOtudent editorial team.
This post is also available in:














