TL;DR. Habit stacking works by attaching a new behaviour to an existing one that reliably happens. That mechanism is sound, but it inherits a hidden requirement: the anchor has to run to completion, in the same context, many times. Two independent bodies of evidence say that requirement is now the binding constraint. Observational research on information workers found the average work block lasts 11 minutes 4 seconds and 57 percent of blocks are interrupted, with a same-day return to the interrupted work taking 25 minutes 26 seconds and passing through 2.26 other topics on the way. A 2024 systematic review of habit formation found that in one study, the same daily stretching routine took a mean of 106 days to automate when done in the morning and 154 days when done in the evening. Same behaviour, same protocol, 48 extra days for the slot alone. Put those together and the practical conclusion changes: in an interrupted workday, anchor choice is not a detail of habit stacking, it is most of the outcome. The correct move is to stop stacking onto tasks and start stacking onto boundaries.
Why habit stacking is being oversold right now
The advice is simple enough that it travels well. Find something you already do without thinking. Attach the new behaviour immediately after it. The old habit becomes the cue for the new one, and you stop relying on memory or motivation.
The mechanism is real. What gets dropped in the retelling is the fine print: the cue has to be stable, the context has to repeat, and the pairing has to happen enough times for automaticity to develop. The literature is unambiguous that this takes a while and varies enormously between people. It is much less discussed that the environment doing the repeating has changed.
Here is the part that makes this a 2026 problem rather than a 2016 one. Enterprise adoption of AI technologies in the European Union was essentially flat between 2021 and 2023, then moved sharply.
Table 1: Enterprise AI adoption, EU27, enterprises with 10 or more employees (verified data)
| Year | Share using at least one AI technology | Change vs previous survey | Relative change |
|---|---|---|---|
| 2021 | 7.65% | – | – |
| 2023 | 8.06% | +0.41 pp | +5.4% |
| 2024 | 13.48% | +5.42 pp | +67.3% |
| 2025 | 19.95% | +6.47 pp | +48.0% |
Source: Eurostat, “Artificial intelligence by size class of enterprise” (isoc_eb_ai), EU27, size class 10 or more employees, percentage of enterprises. Percentage-point and relative-change columns computed by CEOtudent from the published series.
The share of enterprises using AI is 2.61 times its 2021 level, and almost all of that movement happened after 2023. Two consecutive surveys of roughly 50 to 67 percent relative growth is not a trend line, it is a regime change. Nearly every widely circulated habit-stacking guide was written for the work rhythm on the left side of that table.
This matters because agentic tools introduce a category of interruption that older advice never had to model. When you delegate a task to an assistant and it returns three minutes later, you have not saved an interruption, you have scheduled one. The wait is short enough that you start something else and long enough that you have to come back. That is the exact shape of the fragmentation the observational literature describes, except now you are generating it yourself.
What the workday actually looks like
The best public measurements of work fragmentation come from direct observation rather than self-report, which matters because people are poor at estimating their own switching.
In one study, researchers shadowed information workers for three days each, averaging 25 hours 42 minutes of observation per person and more than 700 formal hours in total, coding every event and assigning it to a “working sphere” – a coherent unit of work with its own people, artefacts and goal.
Table 2: Work fragmentation in information work (verified data)
| Measure | Value |
|---|---|
| Average time in a working sphere before switching or being interrupted | 11 min 4 sec (sd 18 min 9 sec) |
| Share of working spheres that were interrupted | 57% |
| Average number of distinct working spheres per person | 11.7 |
| Interrupted work resumed the same day | 77% |
| Time to resume, when resumed the same day | 25 min 26 sec (sd 54 min 48 sec) |
| Other working spheres visited before resuming | 2.26 (sd 2.79) |
| Share of resumptions that were self-initiated | 90.1% |
| Time before self-resumed work was picked back up | 21 min 28 sec (sd 46 min 47 sec) |
| Time before externally resumed work was picked back up | 61 min 37 sec (sd 95 min 11 sec) |
| Length of segments that were interrupted | 12 min 40 sec (sd 14 min 33 sec) |
| Length of segments that were not interrupted | 8 min 58 sec (sd 14 min 43 sec) |
Source: Mark, G., Gonzalez, V. M., and Harris, J. “No task left behind? Examining the nature of fragmented work.” CHI 2005. Figures as published.
Three findings deserve to be pulled out, because each one breaks a different assumption habit stacking makes.
The anchor is more likely to be interrupted than not. Fifty-seven percent is a majority. If your stack is “after I finish reviewing the morning queue, I write my one-paragraph plan,” the reviewing step gets interrupted more often than it completes cleanly.
Interruption does not mean a short detour. Returning to the interrupted work took over 25 minutes on average and passed through more than two unrelated topics first. By the time you are back, the mental context that would have triggered the stacked behaviour is gone.
Ninety percent of resumptions are self-initiated. This is the most uncomfortable number in the table. Most people picture interruption as something done to them. In this data, when work is resumed it is overwhelmingly the person themselves who returns to it – and self-resumed work comes back roughly three times faster than externally resumed work. The interruption environment is substantially internal, which means it does not get fixed by turning off notifications alone. Our notification audit protocol handles the external half; this piece is about designing for the half that remains.
The compensation trap: you speed up, and you pay for it
There is a persistent belief that interruption is survivable because you catch up afterwards. A controlled experiment tested exactly this, comparing an uninterrupted baseline against two interruption conditions – one where the interruption concerned the same topic as the task, one where it concerned a different topic.
The finding was counterintuitive: people completed the interrupted task in less time, with no significant difference in error count. But the workload measures moved sharply in the other direction. What nobody computes from that paper is the exchange rate between the two.
Table 3: The exchange rate of working faster under interruption (CEOtudent derived analysis)
| Measure | Same-context interruption | Different-context interruption |
|---|---|---|
| Time to perform task | -10.80% | -9.53% |
| Reported stress | +36.71% | +31.94% |
| Reported frustration | +40.17% | +37.00% |
| Time pressure | +15.15% | +10.44% |
| Effort invested | +16.21% | +21.26% |
| Mental workload | +8.08% | +14.77% |
| Words written per email | -7.37% | -4.22% |
| Errors per email | -0.52% | -5.15% |
| Stress increase per 1% of time saved | 3.40x | 3.35x |
Derived by CEOtudent from the published means in Mark, G., Gudith, D., and Klocke, U., “The cost of interrupted work: more speed and stress,” CHI 2008, Tables 1 and 3. Baseline condition: 22.77 minutes, stress 6.92, frustration 4.73, time pressure 11.02, effort 9.50, mental workload 10.02, 31.49 words, 1.94 errors. All percentages are changes against that baseline; the final row is the ratio of the stress change to the time change. The workload scale runs from 1 to 20. Time, workload, stress, frustration, time pressure and effort differences were statistically significant in the source; the error-count difference was not.
The bottom row is the finding. Across both interruption types the ratio is almost identical – about 3.4 – which is what you would expect if a single compensation mechanism is doing the work rather than something specific to the interruption’s topic. Every one percent of time you claw back by working faster under interruption costs roughly three and a half percent more subjective stress.
Two honest qualifications. First, the two conditions differ only slightly on time (-10.80% versus -9.53%), so small differences in the ratio should not be over-read; the robust claim is that the ratio is roughly 3.4 in both, not that same-context interruption is worse. Second, part of the apparent speed-up is output shrinkage: emails written under interruption were 4 to 7 percent shorter. The task got finished faster partly because less of it got produced.
For habit stacking, the implication is direct. A stacked behaviour is exactly the kind of low-urgency, self-initiated item that gets cut when you are compensating. It is not that you forget. It is that the compensating version of you does not have room for it.
The number that reframes the whole practice
The habit-formation evidence base was reviewed systematically in 2024, covering 20 studies and 2,601 participants across physical activity, drinking water, vitamin consumption, flossing, diet, sedentary behaviour reduction and microwaving a dishcloth. Eleven of the 20 studies carried a high risk of bias, which the review states plainly and which should temper any confident number here.
Four of those studies measured how long automaticity took.
Table 4: Time to reach habit formation, as reported in the four studies that measured it (verified data)
| Study | Behaviour | Reported time | Reported spread | Note |
|---|---|---|---|---|
| Keller and colleagues | Healthy eating | Median 59 days | Range 4 to 335 days | Only 23% of participants reached the predetermined habit threshold |
| Lally and colleagues (2010) | Eating, water, or exercise | Median 66 days to 95% automaticity | Range 18 to 254 days | The origin of the widely quoted “66 days” |
| Lally and colleagues (earlier) | Healthy eating and self-weighing | Mean 91 days | sd 55 days | Self-reported duration, not measured |
| Fournier and colleagues | Daily stretching, morning | Mean 106 days | – | Same study, same behaviour as the row below |
| Fournier and colleagues | Daily stretching, evening | Mean 154 days | – | Same study, same behaviour as the row above |
Source: Singh, B., Murphy, A., Maher, C., and Smith, A. E. “Time to form a habit: a systematic review and meta-analysis of health behaviour habit formation and its determinants.” Healthcare, 2024, volume 12, issue 23, article 2488. Figures as reported in the review. The review also reports a pooled effect of SMD 0.69 (95% CI 0.49 to 0.88) for the improvement in habit scores from before to after intervention, with substantial heterogeneity.
Now the two rows that do the work. The Fournier study ran the same behaviour under the same protocol and varied only the time of day.
Table 5: The anchor-slot premium (CEOtudent derived analysis)
| Comparison | Morning | Evening | Absolute difference | Ratio |
|---|---|---|---|---|
| Mean days to automate daily stretching | 106 | 154 | 48 days | 1.45x |
Derived by CEOtudent from the Fournier figures reported in Singh and colleagues (2024). This is a within-study, same-behaviour comparison, which is why it is the cleanest number in this article.
Forty-eight days. A 45.3 percent penalty, and the only thing that changed was which slot the behaviour was stacked into. The review’s own synthesis of determinants points the same direction: it reports that morning practices and self-selected habits generally show greater habit strength, and that frequency, timing, type of habit, individual choice, affective judgements, behavioural regulation and preparatory habits all significantly influence habit strength.
Set that next to the fragmentation data and the mechanism is not mysterious. A morning slot sits before the day’s interruption load has accumulated. An evening slot sits after it, downstream of every unresolved working sphere, every 25-minute resumption, and every compensating sprint. The slot is not a preference. It is a measure of how much interference the anchor has to survive.
There is a second, less comfortable reading of Table 4. Keller’s study reports a median of 59 days, but only 23 percent of participants reached the threshold at all. That median describes the people who made it. Every headline habit-formation number is, to some degree, a survivor statistic – it tells you how long it took for the people it worked for, not how likely it is to work.
What interruption does to the repetition count
Habit stacking runs on repetitions in a stable context. The fragmentation data says a majority of blocks do not complete cleanly. That suggests an obvious question nobody seems to have asked: how much does the calendar stretch if a meaningful share of your anchor executions never reach the stack point?
This next table is a model, not a measurement. It is stated in full so you can disagree with it precisely.
Table 6: Projected calendar time under interruption (CEOtudent projection – modelled, not observed)
| Published requirement | Clean repetitions needed | Half-credit projection (1.40x) | Zero-credit projection (2.33x) |
|---|---|---|---|
| Keller median, healthy eating | 59 | 82.5 days | 137.2 days |
| Lally median, 95% automaticity | 66 | 92.3 days | 153.5 days |
| Fournier mean, morning stretching | 106 | 148.3 days | 246.5 days |
| Fournier mean, evening stretching | 154 | 215.4 days | 358.1 days |
CEOtudent projection. Assumptions, stated explicitly: (1) the published day counts are treated as counts of effective repetitions in a stable context; (2) the 57 percent interruption rate from the fragmentation study is applied to the anchor; (3) the zero-credit column assumes an interrupted anchor produces no progress toward automaticity, giving a factor of 1 divided by 0.43, or 2.33; (4) the half-credit column assumes an interrupted anchor still yields half a repetition, giving a factor of 1 divided by 0.715, or 1.40. This is arithmetic on two datasets that were never designed to be combined – the fragmentation study measured office work, the habit studies measured health behaviours. Treat it as a bounding exercise, not a forecast.
The model fails a sanity check at one end, and that failure is informative. The zero-credit projection for evening stretching is 358.1 days, which is longer than the longest individual duration observed anywhere in the reviewed studies (335 days). Since no participant in the underlying literature took that long, the zero-credit assumption is almost certainly too pessimistic – an interrupted attempt evidently does contribute something. The realistic range sits between the two columns, closer to the half-credit side.
The useful conclusion survives the caveat. Even at half credit, an interruption-heavy environment adds roughly 40 percent to the calendar. The Lally median of 66 days becomes about 92. The comfortable “two months” becomes a full quarter. Most people abandon a stack somewhere in the gap between the number they were promised and the number their environment actually charges them.
The Interruption-Resistant Stack: a CEOtudent framework
Standard habit stacking says: find an existing habit, attach the new behaviour after it. Given the evidence above, that instruction is incomplete in one specific way – it treats all anchors as interchangeable when they demonstrably are not. This framework replaces “find an existing habit” with an anchor test, and adds a recovery rule for the majority case where the anchor gets interrupted.
Step 1. Test the anchor for interruptibility, not familiarity.
The standard advice picks anchors by how automatic they already are. Pick instead by how hard they are to interrupt. An anchor qualifies if it meets all four conditions:
- It has a defined end, not just a start. “Checking email” has no end. “Closing the laptop lid” does.
- It takes under two minutes from cue to completion, comfortably inside the 11-minute average block.
- It is not contingent on another person’s response, an approval, or a system finishing a job.
- It happens whether or not the day goes well. Anchors that only occur on good days train the stack to be optional.
Step 2. Prefer boundaries to tasks.
Tasks get interrupted; boundaries do not, because they are transitions rather than work. Sitting down, standing up, the first sip of coffee, unlocking the front door, closing a laptop – these complete in seconds and cannot be half-finished. The fragmentation data shows work blocks fracture. It does not show that transitions fracture, because there is nothing in a transition to fracture.
Step 3. Bias to the earliest viable slot.
This is the 48-day rule. Within the constraints of the behaviour itself, choose the earliest slot in the day where the anchor is genuinely available. The morning-versus-evening gap in Table 5 is the single largest lever in this article and it costs nothing to pull. If a behaviour genuinely cannot happen in the morning, accept the penalty knowingly and budget the extra weeks rather than being surprised by them. Our chronotype self-test is the right input if your biological prime time makes the earliest clock slot the wrong one.
Step 4. Never stack behind an AI assistant’s output.
This is the new rule and it is the one most likely to be violated in 2026. “After the agent finishes drafting, I will review it” is not a stack, it is a dependency on an external completion time you do not control. It fails condition three in Step 1. Stack on your own action – “after I send the prompt, I write the acceptance criteria” – which completes regardless of what the model does or how long it takes. If you are building work routines around agents, our five-layer AI workflow structure covers the layer this sits in, and agent literacy covers what you need to know before delegating at all.
Step 5. Write the recovery branch, not just the plan.
Because 57 percent of blocks are interrupted, a stack that only specifies the happy path fails the majority of the time. Specify the interrupted path in the same sentence: “After I close my laptop I write tomorrow’s first task; if I am pulled away before I close it, I write it standing at the door instead.” The second clause is not a fallback, it is the clause that will run more often than the first.
Step 6. Count anchors, not days.
Do not track a 66-day streak. Track how many times the anchor completed and the stacked behaviour followed. Under a 57 percent interruption rate, 66 calendar days may only contain a fraction of that in effective repetitions, and a streak counter will tell you that you failed when what actually happened is that your environment charged you more days for the same number of repetitions. Counting anchors makes the real progress visible. If you use a tool for this, our assessment of AI habit-tracking assistants covers what they can and cannot verify.
Step 7. Cap the stack at one.
The compensation data explains why. When you are working faster under interruption, subjective effort and stress rise steeply. A stack of four new behaviours is four things competing for the narrow margin that survives that compression. One behaviour, one anchor, until the anchor count says it has taken.
Where this framework should not be used
Being clear about the limits is part of the argument.
This is built for discretionary, self-initiated behaviours in interrupted knowledge work. It is not designed for clinical behaviour change, medication adherence, or addiction recovery, where the supporting structures are different and professional guidance applies.
The evidence base has real limits, stated by the sources themselves. The habit review included 20 studies, 11 of them at high risk of bias, and the durations come from only four of those. The behaviours studied – stretching, flossing, drinking water, diet – are simpler than most knowledge-work routines, and it is plausible that more complex behaviours behave differently. The fragmentation study observed a specific population of information workers at one organisation, and its numbers should be read as a well-measured example rather than a global constant. The interruption experiment was a laboratory task, not a workplace.
And Table 6 is a projection built by joining datasets that were never intended to meet. It is included because the bounding exercise is genuinely useful and because the alternative – quoting “66 days” as if the environment were free – is worse. It is not evidence of a measured effect.
FAQ
Is habit stacking actually broken?
No. The mechanism holds. What is broken is the assumption that any familiar behaviour makes an equally good anchor. The 48-day morning-versus-evening gap in a single controlled comparison shows that anchor selection carries more weight than the technique itself.
Why do so many sources say 21 days?
Because it is easier to repeat than to check. Every measured figure in Table 4 is far longer: medians of 59 and 66 days, means of 91, 106 and 154 days, and observed individual durations stretching to 335 days. No study in the 2024 review supports 21 days as a typical figure.
If 66 days is a median, what happens to everyone else?
That is the right question and it is rarely asked. In the Lally study the observed range ran from 18 to 254 days. In the Keller study only 23 percent of participants reached the habit threshold at all, meaning the reported median describes the minority for whom it worked. Plan for a distribution, not a deadline.
Does an AI assistant help or hurt habit formation?
Both, and it depends entirely on where you put it. Used as a prompt or a log after your own action, it is neutral to helpful. Used as the anchor itself – waiting for output before your behaviour can start – it introduces a dependency on a completion time you do not control, which is exactly the structure that fails under the fragmentation data.
Does turning off notifications solve this?
Only partly, and less than most people assume. In the observational data, 90.1 percent of resumptions were self-initiated. Silencing external interruptions addresses a real but minority share of the switching, which is why Steps 2, 4 and 5 target the structure of the routine rather than the device.
How do I know whether the stack is working?
Count anchor completions followed by the behaviour, and look for the point at which the behaviour starts feeling unremarkable rather than effortful. Automaticity, not streak length, is what the research measures. Track the ratio – behaviour completions divided by anchor completions – and expect it to be the thing that moves first.
Is the evening penalty universal?
It should not be treated that way. It comes from one study of one behaviour, reported in the 2024 review, and it is one comparison rather than a meta-analytic finding. It is highlighted here because it is a rare within-study, same-behaviour test of slot choice, which makes it unusually clean – not because it has been replicated broadly.
Sources
Mark, G., Gonzalez, V. M., and Harris, J. No task left behind? Examining the nature of fragmented work. Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, CHI 2005, Portland, Oregon, pages 321 to 330.
Mark, G., Gudith, D., and Klocke, U. The cost of interrupted work: more speed and stress. Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, CHI 2008.
Singh, B., Murphy, A., Maher, C., and Smith, A. E. Time to form a habit: a systematic review and meta-analysis of health behaviour habit formation and its determinants. Healthcare, 2024, volume 12, issue 23, article 2488. University of South Australia.
Lally, P., van Jaarsveld, C. H. M., Potts, H. W. W., and Wardle, J. How are habits formed: modelling habit formation in the real world. European Journal of Social Psychology, 2010, volume 40, issue 6, pages 998 to 1009. Figures cited here as reported in Singh and colleagues (2024).
Eurostat. Artificial intelligence by size class of enterprise, dataset isoc_eb_ai. European Union, EU27, enterprises with 10 or more employees, survey years 2021, 2023, 2024 and 2025.
Derived figures in Tables 1, 3 and 5, and the projections in Table 6, were computed by CEOtudent from the published values cited above and recomputed independently before publication. Table 6 is explicitly a model and is labelled as such.
This content was compiled with the support of AI following in-depth research, then written and prepared for publication by the CEOtudent editorial team.
This post is also available in:














