TL;DR: Decision fatigue is two different claims wearing one name, and only one of them survived the last decade of scrutiny. The laboratory claim, that self-control runs on a depletable resource, failed two of the largest coordinated replication efforts in psychology: Hagger and colleagues (2016) pooled 23 laboratories and 2,141 participants and found d = 0.04, 95 percent CI [-0.07, 0.15]; Vohs and colleagues (2021) pooled 36 laboratories and 3,531 participants and found d = 0.06, with Bayesian analysis putting the data four times more likely under the null than under the hypothesis. The glucose mechanism underneath it does not survive a metabolic audit either. Meanwhile the field claim, that decision quality drifts across a working session, keeps showing up in administrative data: antibiotic prescribing odds rise from 1.01 in hour two to 1.26 in hour four across 21,867 clinic visits, and Danish national test data shows performance falling 0.9 percent of a standard deviation per hour with a 20-30 minute break recovering 1.7 percent. Below: an original evidence ledger grading each popular claim, a derived conversion showing what a single break is actually worth in hours, and a day structure built only on the surviving findings.
Almost every article about decision fatigue opens the same way, with a judge, a parole hearing, and a lunch break. It is a good story. It is also the single most contested finding in the literature, and the writers repeating it have almost never read the three papers that followed it.
That is the problem with this topic. Decision fatigue is not wrong. It is over-claimed in one direction and under-used in another, and the gap between those two things is where most people’s day structure goes wrong. If you build your schedule on the collapsed version of the theory, you will spend your energy on the wrong intervention: eating something, willing yourself harder, saving your big call for a morning that is no longer fresh by the time you get to it.
This piece does the boring thing. It separates what failed from what held, and then builds the schedule from the survivors only.
The claim that collapsed
The willpower-as-fuel model, formally called ego depletion, held that self-control draws on a limited common resource, so exercising restraint on one task leaves less available for the next. It produced a decade of confident advice and a meta-analysis in 2010 reporting a medium effect of d = 0.62.
Then people started checking.
Carter and McCullough (2014) applied publication-bias corrections to that same dataset and watched the effect move depending on the correction used, landing at d = 0.48 with trim-and-fill and at d = 0.25 or even d = -0.10 with regression-based procedures. Their conclusion was not that the effect was zero but that the evidence could not distinguish it from zero.
Two coordinated multi-laboratory tests then ran the experiment properly, preregistered, with the analysis plan fixed in advance. Hagger and colleagues (2016) reported d = 0.04 with a confidence interval spanning zero across 23 laboratories. Vohs and colleagues (2021), using a different and deliberately more forgiving design intended to give the effect its best shot, reported d = 0.06 across 36 laboratories and found the data roughly four times more likely under the null hypothesis than under an informed prior.
The metabolic story fared no better. Kurzban (2010) worked out the actual energy cost of a self-control task and found it implausibly small, on the order of a fraction of a calorie, because the brain metabolises glucose at broadly similar rates across very different cognitive tasks. The brain does not run out of fuel between your second and third meeting.
None of this means willpower is imaginary. It means the specific mechanism people were sold, a tank that empties and refills with sugar, is not the mechanism.
The claim that survived
Here is where most coverage stops, and it stops too early. The lab paradigm failing does not settle the question of whether decision quality drifts across a real working day, because those are different measurements. Field studies do not manipulate self-control with a word puzzle. They observe thousands of consequential decisions in sequence.
Linder and colleagues (2014), publishing in JAMA Internal Medicine, examined 21,867 acute respiratory infection visits to 204 clinicians across 23 primary care practices. Clinicians worked in four-hour morning and afternoon sessions. Relative to the first hour of a session, the adjusted odds of prescribing an antibiotic were 1.01 in hour two (95 percent CI 0.91-1.13), 1.14 in hour three (1.02-1.27), and 1.26 in hour four (1.13-1.41), with p less than .001 for the linear trend. Overall, 44 percent of visits ended in a prescription, and the case mix did not shift meaningfully across the session.
Sievertsen, Gino and Piovesan (2016), using Danish administrative records covering all children in public schools across four school years, found test performance falling by 0.9 percent of a standard deviation for each hour later in the day (95 percent CI 0.7-1.0 percent), and a 20-30 minute break raising average performance by 1.7 percent of a standard deviation (95 percent CI 1.2-2.2 percent).
Note what these two studies have in common and what they do not share with the lab paradigm. They measure a sequence position inside a bounded working block. They observe real decisions with real stakes. And crucially, in the Linder data the drift resets: the afternoon session starts over at hour one, and the pattern repeats. That detail matters enormously for scheduling, and we will come back to it.
The evidence ledger
The following table is a CEOtudent editorial framework. It grades the popular claims of decision fatigue against the strongest published test of each. The evidence tiers are our own classification, applied consistently: Collapsed means large preregistered tests found effects indistinguishable from zero; Contested means the finding stands but a documented confound offers an alternative explanation; Holds means the finding comes from large field or administrative datasets and has not been overturned.
| Popular claim | Strongest test of it | Reported result | Tier | What you should do with it |
|---|---|---|---|---|
| Self-control is a depletable resource | Hagger et al. 2016, 23 labs, N = 2,141 | d = 0.04, CI [-0.07, 0.15] | Collapsed | Stop budgeting willpower; budget decisions |
| Depletion holds under a fairer design | Vohs et al. 2021, 36 labs, N = 3,531 | d = 0.06, data 4x likelier under null | Collapsed | Do not expect a second wind from resting willpower |
| Glucose refuels self-control | Kurzban 2010, metabolic audit | Cost of task well under 0.2 calories | Collapsed | Snacks are not a decision-quality intervention |
| Original evidence base was solid | Carter and McCullough 2014, bias correction | d falls from 0.62 to between -0.10 and 0.48 | Collapsed | Treat pre-2015 willpower advice as unverified |
| Judges rule harshly before lunch | Danziger et al. 2011; Glöckner 2016 simulation | Effect real in data; scheduling artifact reproduces it | Contested | Stop citing it as the proof; it is the weakest link |
| Decision quality drifts within a session | Linder et al. 2014, N = 21,867 visits | OR 1.26 by hour four, p < .001 trend | Holds | Cap session length and order by stakes |
| Breaks restore performance | Sievertsen et al. 2016, Danish national tests | +1.7 percent of SD from a 20-30 min break | Holds | Schedule breaks as infrastructure, not reward |
| Deciding in advance beats deciding in the moment | Gollwitzer and Sheeran 2006, k = 94, N over 8,000 | d = 0.65 on goal attainment | Holds | Convert recurring choices into if-then rules |
The single most useful line in that table is the last one. The intervention with the largest surviving effect size in this entire area is not resting, eating, or scheduling. It is removing the decision from the day altogether by pre-committing to a rule.
The hungry judges problem, briefly
Danziger, Levav and Avnaim-Pesso (2011) reported that favourable parole rulings fell from roughly 65 percent at the start of a session toward near zero by its end, then reset after a food break. It became the emblem of decision fatigue.
Two objections followed. Weinshall-Margel and Shapard (2011) established from interviews with court personnel that case ordering was not random: the board completed all cases from one prison before breaking, and unrepresented prisoners, who are less likely to be granted parole, typically appeared late in sessions. Glöckner (2016) then showed by simulation that a purely mechanical scheduling artifact can reproduce the pattern, because favourable rulings take longer to process, so a panel managing its remaining time will systematically place quicker unfavourable cases at the end of a block.
The honest summary is that the pattern in the data is real and the causal story is not established. This is why the ledger above rates it Contested rather than Collapsed or Holds, and why building your day around it is building on the weakest available beam when two stronger ones are sitting right there.
What a break is actually worth
The Danish dataset lets us compute something that is not stated in the paper and is directly useful for scheduling. If performance declines by 0.9 percent of a standard deviation per hour, and a 20-30 minute break recovers 1.7 percent of a standard deviation, then the break buys back roughly 1.9 hours of accumulated decline.
Working the confidence intervals through gives the plausible range. At the pessimistic end, the smallest break effect (1.2 percent) against the steepest hourly decline (1.0 percent) gives 1.2 hours. At the optimistic end, the largest break effect (2.2 percent) against the shallowest decline (0.7 percent) gives 3.1 hours.
| Quantity | Point estimate | Range from reported CIs |
|---|---|---|
| Decline per hour on task | 0.9 percent of SD | 0.7-1.0 percent |
| Recovery from a 20-30 min break | 1.7 percent of SD | 1.2-2.2 percent |
| Hours of decline a break offsets (derived) | About 1.9 hours | About 1.2-3.1 hours |
| Implied maximum useful block length | About 2 hours | About 1-3 hours |
This is a CEOtudent-derived calculation from the published Danish figures, not a result reported by the authors, and it inherits every limitation of the original study, including that it measures schoolchildren on standardised tests rather than adults on knowledge work. Treat it as an order-of-magnitude guide, not a law. What it gives you is a defensible answer to a question people usually answer by vibe: a break is worth about two hours of drift, so a block much longer than two hours is spending capacity you will not get back with one pause.
The reframe that changes your calendar
Here is the thing almost all decision fatigue advice gets wrong. It tells you to make important decisions “in the morning.”
The Linder data says something different and more useful. The drift is measured against session position, not clock time. Hour one of the afternoon session behaves like hour one of the morning session. The reset is not sunrise. The reset is the boundary.
That single distinction rewrites the practical advice:
- The question is not “is it early in the day” but “is it early in this block.”
- A 2 pm decision made in the first hour after a real break is in better shape than an 11 am decision made in the fourth hour of a session that started at 8.
- Someone who never takes a genuine break has one enormous session, and by mid-afternoon they are permanently in hour six of it.
This also explains why the popular advice fails for people who are not morning people, who have caregiving mornings, or who work across time zones. They were being told to optimise a variable that was never the operative one.
Structuring the day around what holds
Five rules, each traceable to a finding in the ledger rather than to the collapsed theory.
1. Cut the number of decisions before you optimise their placement. Gollwitzer and Sheeran’s d = 0.65 across 94 tests is the largest surviving effect in this space. Any recurring choice you make more than weekly should become an if-then rule with the condition written explicitly: if it is Monday morning, then I review the pipeline before opening email. This is not the same as a habit or a preference. It is a rule with a named trigger, which is why it outperforms intention.
2. Bound your sessions at roughly two hours and make the boundary real. The derived break value gives about 1.9 hours as the point where a single break stops covering the accumulated drift. A boundary is only a boundary if something changes: standing up, leaving the room, changing the physical context. A “break” spent scrolling in the same chair on the same screen is a continuation of the session with worse inputs.
3. Put consequential decisions in the first hour of a block, whichever block that is. From the Linder odds pattern, position one is the good seat. Note that the hour-two odds ratio was 1.01 with a confidence interval crossing one, meaning the meaningful degradation shows up in hours three and four, not immediately. You have roughly two clean hours per block, which is exactly consistent with the derived break arithmetic.
4. Do not treat food as the intervention. Kurzban’s metabolic audit removed the mechanism. Eat because you are hungry and because hunger is distracting, not because you believe you are refuelling a control system. If a meal helps, it is almost certainly acting as a genuine break in context and attention rather than as fuel.
5. Log the misses rather than trusting the feeling. The fatigue signal is unreliable, which is the whole lesson of the last ten years. What is reliable is the record: which decisions you later reversed, which emails you regret sending, which approvals you gave without reading. If those cluster in the back half of your blocks, you have your own version of the Linder curve, measured on the only dataset that matters to you. Our Energy Auditing method covers the tracking mechanics, and the Cognitive Load Budget piece covers how to read the resulting pattern.
The CEO and the student in this
The CEO move here is structural. You do not solve a drift problem by trying harder inside the drift. You change the shape of the container: fewer decisions reach you, the ones that do arrive early in a bounded block, and the boundaries are enforced by the calendar rather than by how you feel. That is designing your energy infrastructure rather than negotiating with it, and it is the same logic as the personal decision stack: put the judgment into the system once, so the system carries it every day.
The student move is harder and rarer. Ego depletion was taught in undergraduate courses, cited tens of thousands of times, and built into books and companies. Then it did not replicate, twice, at scale, under preregistration. The people who updated look inconsistent to anyone who only heard the first version. Updating anyway is the actual skill, and it is the one that compounds: a practice built on the 2011 evidence and never revised is now, quietly, a practice built on something the field has largely abandoned.
Most people will keep repeating the judges. Do the other thing.
FAQ
Is decision fatigue real or not?
Both answers are partly right, which is why the topic is such a mess. The laboratory theory that self-control depletes like a resource is not supported: 23 labs found d = 0.04 and 36 labs found d = 0.06. The field observation that decision quality drifts across a working session is supported by large administrative datasets, including 21,867 clinic visits and Danish national test records. The pattern is real; the fuel-tank explanation for it is not.
Should I stop citing the hungry judges study?
As proof, yes. The pattern in that dataset is genuine, but case ordering was not random, unrepresented prisoners appeared late in sessions, and a simulation showed a scheduling artifact can reproduce the effect on its own. There are two much stronger findings available. Use those instead.
How long should a work block be?
Deriving from the Danish figures, a 20-30 minute break offsets roughly 1.9 hours of accumulated decline, with a plausible range of about 1.2-3.1 hours. That points to blocks of around two hours as a reasonable default. The Linder data agrees from the other direction: degradation was not statistically distinguishable in hour two but was clear by hours three and four.
Does eating restore decision quality?
There is no supported metabolic mechanism. Kurzban’s audit put the caloric cost of a self-control task well under 0.2 calories, and brain glucose consumption does not vary much across cognitive tasks. Eat for the ordinary reasons. If it helps your decisions, the likely active ingredient is the break itself.
Is it better to decide in the morning?
Less than you have been told. The strongest field evidence measures position within a session rather than time on the clock, and the afternoon session in the Linder data starts fresh. The first hour after a genuine break beats the fourth hour of a long morning.
What is the single highest-value change?
Converting recurring decisions into if-then rules. That intervention has the largest surviving effect size in the area, d = 0.65 across 94 independent tests with more than 8,000 participants, and unlike scheduling tweaks it removes the decision permanently rather than relocating it.
Does any of this apply to AI-assisted work?
It applies more, not less. Tools that generate options faster increase the number of choices reaching you per hour, which loads the exact variable the field evidence implicates. The counterweight is deciding in advance which classes of choice you will not personally make, which is the question covered in what to delegate and what never to automate.
Sources
- Hagger and colleagues, A Multilab Preregistered Replication of the Ego-Depletion Effect, Perspectives on Psychological Science, 2016, 23 laboratories, 2,141 participants
- Vohs and colleagues, A Multisite Preregistered Paradigmatic Test of the Ego-Depletion Effect, Psychological Science, 2021, 36 laboratories, 3,531 participants
- Carter and McCullough, Publication bias and the limited strength model of self-control, Frontiers in Psychology, 2014
- Kurzban, Does the Brain Consume Additional Glucose during Self-Control Tasks, Evolutionary Psychology, 2010
- Linder and colleagues, Time of Day and the Decision to Prescribe Antibiotics, JAMA Internal Medicine, 2014, 21,867 visits across 23 practices
- Sievertsen, Gino and Piovesan, Cognitive fatigue influences students’ performance on standardized tests, Proceedings of the National Academy of Sciences, 2016
- Danziger, Levav and Avnaim-Pesso, Extraneous factors in judicial decisions, Proceedings of the National Academy of Sciences, 2011
- Weinshall-Margel and Shapard, Overlooked factors in the analysis of parole decisions, Proceedings of the National Academy of Sciences, 2011
- Glöckner, The irrational hungry judge effect revisited, Judgment and Decision Making, 2016
- Gollwitzer and Sheeran, Implementation intentions and goal achievement: A meta-analysis of effects and processes, Advances in Experimental Social Psychology, 2006
This content was compiled with the support of AI following in-depth research, then written and prepared for publication by the CEOtudent editorial team.
This post is also available in:
















