TL;DR: A decision journal is not a diary. It is a measurement instrument, and it only works if it is built to defeat a specific, replicated memory failure. Fischhoff established in 1975 that outcome knowledge silently rewrites what you believe you expected, by a mean of 9.2 percentage points across his reported outcomes, which is why the entry must be written before the result exists and must contain a number. Baron and Hershey established in 1988 that people judge an identical decision as better when the outcome happened to be good, and keep doing so after being explicitly instructed not to, which is why the journal has to separate decision quality from outcome quality as two scored fields. Mellers and colleagues, reporting on a two-year geopolitical forecasting tournament covering 199 questions, found that a probability training module taking about 45 minutes improved accuracy across periods 8 to 10 months long, and that the tournament’s elite forecasters made an average of 7.8 forecasts per question while independent forecasters made 1.4. Mitchell, Russo and Pennington found that imagining an outcome has already happened raises your ability to correctly identify reasons for it by 30 percent. Put together, those four results specify the instrument exactly: seven fields, a calibrated probability, a pre-written failure story, and a review loop that fires on a trigger rather than a mood. Below: the template, a derived table of what calibration training actually bought at each stage of a forecast, the scoring rules, and the honest limits.
There is a version of self-improvement that feels like progress and produces none. You write down what you decided. Months later you read it back, notice you were right about some things, feel encouraged, and change nothing.
The reason this fails is not laziness. It is that the human memory of a prediction is not a stored record. It is reconstructed at the moment you retrieve it, and the reconstruction is contaminated by everything you have learned since.
A decision journal that does not account for that is a diary. A decision journal that does account for it is one of the few instruments that can actually move judgment.
The four findings the template has to satisfy
Everything structural about the format below follows from four results. They are worth stating precisely, because each one dictates a specific field.
Your memory of what you expected moves after you learn the answer. Baruch Fischhoff’s 1975 paper, published under the title Hindsight is not equal to foresight, gave people brief historical and clinical vignettes with several plausible outcomes and asked for probabilities. Some were told which outcome occurred. Across the reported outcomes, the mean difference between foresight and hindsight probabilities was 9.2 percentage points. People asked to ignore the outcome could not. The effect became one of the most replicated results in the field; Christensen-Szalanski and Willham’s meta-analysis pooled 122 studies of it.
The design consequence: the entry has to be timestamped before the outcome exists, and it has to contain a number, because a remembered feeling is exactly the thing that drifts.
Your evaluation of a decision moves with the result, even when you know it should not. Baron and Hershey’s 1988 study in the Journal of Personality and Social Psychology showed that an identical decision was rated as better, more important and more normative when the outcome was positive. Subjects who were asked agreed that outcomes should not affect the evaluation. They were affected anyway.
The design consequence: decision quality and outcome quality must be two separate scored fields, filled in at two separate times. If one field carries both, the good result will quietly upgrade the reasoning that produced it.
A short, one-time calibration exercise has a long half-life. Mellers and colleagues reported in Psychological Science on a two-year tournament in which forecasters answered 199 geopolitical questions, open for an average of 102 days each. Participants were randomly assigned to probability training, scenario training or no training. Probability training beat scenario training, which beat no training. The module took about 45 minutes to complete, and its benefit persisted across two periods each roughly 8 to 10 months long.
The design consequence: it is worth spending one session learning to state probabilities properly, and the journal should keep a running calibration score rather than an impression.
Imagining the failure before it happens improves your causal reasoning. Mitchell, Russo and Pennington, in a 1989 paper in the Journal of Behavioral Decision Making, found that prospective hindsight, describing a future event as though it had already occurred, increased the ability to correctly identify reasons for outcomes by 30 percent. This is the mechanism underneath the premortem.
The design consequence: there is a required field in which you write, in past tense, the story of how this went wrong.
What 45 minutes of calibration training actually bought
The tournament data lets us do something the coverage of it rarely does: compute the size of the training effect at each stage of a forecast, rather than repeating the headline. Brier scores measure probabilistic accuracy and lower is better, so an improvement is a reduction. The table below derives the relative improvement from probability training against the matched no-training condition, cell by cell, from the published score table.
Table 1. Relative accuracy gain from a 45-minute probability training module, by forecaster type and forecast stage, CEOtudent editorial framework, derived from published Brier scores
| Year | Forecaster type | Stage | Brier, no training | Brier, probability training | Relative improvement |
|---|---|---|---|---|---|
| 1 | Individual | First week | 0.44 | 0.40 | 9.1% better |
| 1 | Individual | Middle 2 weeks | 0.40 | 0.36 | 10.0% better |
| 1 | Individual | Last week | 0.31 | 0.29 | 6.5% better |
| 1 | Crowd-belief | First week | 0.42 | 0.36 | 14.3% better |
| 1 | Crowd-belief | Middle 2 weeks | 0.39 | 0.34 | 12.8% better |
| 1 | Crowd-belief | Last week | 0.30 | 0.23 | 23.3% better |
| 1 | Team | First week | 0.42 | 0.35 | 16.7% better |
| 1 | Team | Middle 2 weeks | 0.33 | 0.30 | 9.1% better |
| 1 | Team | Last week | 0.22 | 0.19 | 13.6% better |
| 2 | Individual | First week | 0.46 | 0.42 | 8.7% better |
| 2 | Individual | Middle 2 weeks | 0.39 | 0.36 | 7.7% better |
| 2 | Individual | Last week | 0.26 | 0.24 | 7.7% better |
| 2 | Team | First week | 0.38 | 0.40 | 5.3% worse |
| 2 | Team | Middle 2 weeks | 0.32 | 0.28 | 12.5% better |
| 2 | Team | Last week | 0.16 | 0.16 | no change |
Three things in that table matter more than the averages.
The first is that the gain is real but modest for a lone individual: between 6.5 and 10 percent across every individual cell in both years. If you are journalling alone, that is the honest size of what a calibration habit buys you directly.
The second is that two cells go the wrong way. In Year 2, trained team forecasters were slightly worse in the first week and identical in the last. A table that hid those would be a sales document. The effect is real and it is not uniform.
The third is the largest single number in the table, 23.3 percent, and it is not in the individual rows. It appears where training combined with seeing what others believed, late in a question’s life, when information had accumulated. Training pays most when it has something to work on.
The finding that actually explains superforecasters
The tournament’s top performers were promoted into elite teams in Year 2, and their scores were not close to anyone else’s: 0.25 in the first week of a question, 0.19 in the middle, 0.07 in the last. Against untrained individual forecasters in the same year, that is 45.7 percent better early, 51.3 percent better in the middle, and 73.1 percent better at the end.
The behavioural difference underneath that is startlingly simple, and it is the single most important thing in this article. Independent forecasters made an average of 1.4 forecasts per question. Crowd-belief forecasters made 1.3. Regular teams in Year 2 made 1.6. The superforecasters made 7.8.
That is a revisit rate roughly 5.6 times higher than independent forecasters. Not a better first guess. More returns to the same open question, each one an update against new information.
This is the finding that should decide how you build your journal. Almost every decision journal template treats the entry as the product. The evidence says the entry is the cheap part. The performance lives in the loop that brings you back to an open decision while it is still open, and most templates have no such loop at all.
Table 2. Revisit rate by forecaster group, and the accuracy that came with it, CEOtudent editorial framework, derived from published tournament figures
| Group | Average forecasts per question | Revisit rate vs independent | Year 2 Brier, last week |
|---|---|---|---|
| Crowd-belief forecasters | 1.3 | 0.9x | not reported separately in Year 2 |
| Independent forecasters | 1.4 | 1.0x baseline | 0.26 |
| Regular teams, Year 2 | 1.6 | 1.1x | 0.16 |
| Superforecasters | 7.8 | 5.6x | 0.07 |
The causal direction is not established by this comparison; the superforecasters were selected for accuracy first and then given elite teams, so their revisit rate and their accuracy share a common cause in ability and engagement. What the numbers establish is the association, which is strong enough to make the revisit loop the part of the instrument you should refuse to skip.
The template: seven required fields
Each field exists because one of the findings above demands it. Nothing here is decorative, and the fields you will be tempted to drop are the two that do the work.
1. Decision and date. One sentence, in the form “I am choosing X over Y.” If you cannot name the alternative you rejected, you have not made a decision, you have had a preference. Timestamp it. The timestamp is what makes the entry admissible later.
2. The situation as it appears now. Three to five sentences. What you know, what you do not, what pressure you are under, how much time you have. Written now, while it is still incomplete, because the version you will remember is the tidy one.
3. The probability. A number, not a word. “I think this works” is unfalsifiable. “I think there is a 65 percent chance this works, judged by the following observable outcome” is a claim you can be scored on. Define the resolution condition in the same field: what specifically will have happened, and by when, for this to count as having worked.
This is the field people skip, and skipping it collapses the journal back into a diary. The 9.2 point drift Fischhoff measured is exactly the gap a number closes.
4. The premortem. In past tense, as though it is a year later and it failed: “This did not work because…” Two or three sentences. Mitchell, Russo and Pennington’s 30 percent improvement in identifying reasons comes from the tense, not the effort. Writing “it might fail if” produces noticeably thinner reasoning than writing “it failed because.”
5. What would change my mind. One or two specific, observable signals that should make you reverse or revise. This field is the trigger for the revisit loop. Without it, you have no scheduled reason to come back, and you will come back at 1.4 times per question like everyone else.
6. Decision quality score, filled now. One to five, on the reasoning alone, before any outcome exists. Did you consider the alternative, look for disconfirming evidence, state a number, name the reversal condition?
7. Outcome quality score, filled at resolution. One to five, on what actually happened, in a separate session. Baron and Hershey’s result is the whole reason these are two fields. When you review, the interesting rows are the mismatches: high decision quality with a bad outcome, which is variance and should not be corrected; and low decision quality with a good outcome, which is the most dangerous row in the journal because nothing about it feels like a problem.
The review protocol: four triggers, not a schedule
Monthly reviews fail for a predictable reason. By the time a month has passed, most entries are still open and the ones that resolved did so at a random point you were not present for. The review has to fire on events instead.
Trigger one: resolution. The moment an outcome is known, before you have talked to anyone about it, open the entry, read your stated probability without editing it, record the outcome, and fill field seven. This is the only moment at which you can compare a preserved prediction against a result. Fifteen minutes.
Trigger two: the change-of-mind signal. When something in field five occurs, the entry gets reopened whether or not the decision is resolved. This is the mechanism that gets you from 1.4 revisits to something closer to the number that correlates with accuracy.
Trigger three: batch calibration, once a quarter. Take every entry that has resolved. Group them by the probability you stated: everything in the 60 to 70 percent band, everything in the 70 to 80 band, and so on. In each band, count what fraction actually happened. If you said 70 percent eleven times and eight of them happened, you are calibrated in that band. If you said 90 percent ten times and six happened, you are overconfident, and you now know it as a number rather than a suspicion. Twenty entries is enough to see a pattern; ten is not.
Trigger four: the mismatch audit, once a year. Pull only the rows where decision quality and outcome quality disagree by two points or more. These are the entries that teach. Everything else mostly confirms what you already believed.
What this instrument cannot do
Being straight about the limits is part of what makes the rest usable.
The tournament evidence comes from geopolitical questions with clean resolution criteria and defined deadlines. Most consequential personal and professional decisions have neither. Whether taking a particular job was correct may not resolve in a scoreable way for years, if ever, and the counterfactual is permanently unavailable. Calibration training transfers to the class of decisions you can actually score, which is a real but bounded subset.
The 9.2 point hindsight effect and the outcome bias result are laboratory findings with vignettes. They establish the mechanism robustly. They do not establish that keeping a journal removes the bias from your working life, and no study cited here measured that.
The superforecaster revisit figure is an association within a selected group, not a demonstration that revisiting more will make you more accurate. Treat 7.8 as evidence about where the leverage plausibly sits, not as a target to hit mechanically.
And the honest base rate for this habit is that most people abandon it. The failure mode is almost always the same: the template had too many fields, the review had no trigger, and the probability field got skipped because it felt pretentious. Seven fields, four triggers, and the number is not optional.
Where this fits
A decision journal is the measurement layer under everything else in decision practice. It is what tells you whether the frameworks you have adopted are actually improving your calls, or whether you simply feel more organised. If you are assembling the wider system, the personal decision stack is the architecture this instrument measures, thinking in bets is where the probability field comes from, and how good judgment actually develops is the research on what the feedback loop is doing to you over years. On the question of which decisions belong in the journal at all, decision fatigue and day structure and the boundaries of delegating a decision to AI draw the two useful lines.
Frequently asked questions
How many decisions should go in the journal?
Fewer than you think. The threshold is: would I want to know, a year from now, what I actually expected here? For most people that is one to three entries a week. A journal with eighty trivial entries and no resolved ones teaches nothing, and the quarterly calibration pass needs roughly twenty resolved entries in a band to say anything.
Does it have to be written, or can it be a conversation with an AI model?
The evidence speaks to the fields, not the medium. A model can help you generate the premortem and can hold you to stating a number, both genuinely useful. What it cannot do is preserve the entry against your own memory, which is the entire point, so whatever you use has to produce a timestamped record you do not edit afterwards.
What if I do not know the probability?
That is the normal condition and it is not an obstacle. Give the number anyway, then let the quarterly calibration pass tell you how wrong your numbers are as a class. Being poorly calibrated and knowing the size of the error is a strictly better position than having no number, and the tournament result suggests one focused session on how to state probabilities is enough to start.
Is the 45-minute training something I can replicate?
Partly. The published module taught reference-class thinking, corrected common biases, and offered heuristics such as averaging multiple estimates. You can practise those directly. What you cannot replicate alone is the tournament’s feedback density, and the table above suggests that seeing what others believed is where the largest gains appeared.
Why score decision quality before the outcome instead of at review?
Because after the outcome you are no longer able to. That is precisely what Baron and Hershey demonstrated: the instruction to ignore the result does not work. The pre-outcome score is the only uncontaminated reading you will ever get of your own reasoning.
Sources
- Fischhoff, Hindsight is not equal to foresight: The effect of outcome knowledge on judgment under uncertainty, Journal of Experimental Psychology: Human Perception and Performance, volume 1, issue 3, 1975, pages 288 to 299
- Christensen-Szalanski and Willham, The hindsight bias: A meta-analysis, Organizational Behavior and Human Decision Processes, volume 48, 1991
- Baron and Hershey, Outcome bias in decision evaluation, Journal of Personality and Social Psychology, volume 54, issue 4, 1988, pages 569 to 579
- Mellers, Ungar, Baron, Ramos, Gurcay, Fincher, Scott, Moore, Atanasov, Swift, Murray, Stone and Tetlock, Psychological Strategies for Winning a Geopolitical Forecasting Tournament, Psychological Science, volume 25, issue 5, 2014
- Mitchell, Russo and Pennington, Back to the future: Temporal perspective in the explanation of events, Journal of Behavioral Decision Making, volume 2, issue 1, 1989
- Tetlock, Mellers, Rohrbaugh and Chen, Forecasting Tournaments: Tools for Increasing Transparency and Improving the Quality of Debate, Current Directions in Psychological Science, volume 23, issue 4, 2014
- Brier, Verification of forecasts expressed in terms of probability, Monthly Weather Review, volume 78, 1950
This content was compiled with the support of AI following in-depth research, then written and prepared for publication by the CEOtudent editorial team.
This post is also available in:















