GelişimStrateji
0

Foresight Methods Ranked: Which Futures-Thinking Techniques Actually Work for Individuals

A woman standing at a sunlit window with an open notebook, looking ahead and thinking about what comes next

TL;DR. Most futures-thinking advice is sold on plausibility rather than measured accuracy, and the two do not correlate well. Working only from published results we could open and recompute, four findings stand out. First, the intervention with the largest measured effect on individual forecasting accuracy is not a technique at all: it is selection by track record, worth a 73.1% reduction in error in the strongest published condition and a 20 to 23% reduction in a separate 2024 study. Second, the best technique changes with timing: in the first week of a forecasting question, probability training beat group discussion; by the final week, discussion was 4.5 times the stronger lever. Third, scenario training, the family that most corporate foresight work belongs to, delivered exactly 0.0% improvement for individuals in the middle period of a question and made team forecasts 9.1% worse in the final week. Fourth, and most uncomfortable, a 2025 re-analysis published in the same journal as the original found that once uncontrolled method variance was modelled, the training and teaming effects were substantially reduced, eliminated, or in one case reversed. What is left is a short list, and it is not the list the foresight industry sells.

Why this question is answerable at all

Futures work has a measurement problem it rarely admits. Most techniques are evaluated on whether participants found the workshop valuable, not on whether the resulting forecasts were closer to what happened. Those are different questions and they can have opposite answers.

There is one large exception. Between 2011 and 2015 the Intelligence Advanced Research Projects Activity ran the Aggregative Contingent Estimation tournament, in which competing research groups recruited thousands of ordinary volunteers, randomly assigned them to conditions, and scored every forecast against real geopolitical outcomes using a proper scoring rule. That design is what makes ranking possible: the techniques were not compared on satisfaction, they were compared on error.

The results were published by Mellers and colleagues in Psychological Science in 2014, and re-analysed by Hauenstein and colleagues in the same journal in 2025. Everything below is built on those two papers plus three other bodies of published evidence. If you keep a record of your own calls, this pairs directly with the decision journal protocol and with thinking in bets.

The raw evidence

Brier scores measure the accuracy of probabilistic forecasts. Lower is better. A forecaster who says 90% and is right scores 0.02 on the scale used in the paper; one who says 90% and is wrong scores 1.62.

Verified data. Average Brier score by condition and by period in a question’s life (Mellers et al., Psychological Science, 2014, Table 1). Only questions open at least a month are included.

Condition First week Middle 2 weeks Last week
Year 1
Individual, no training 0.44 0.40 0.31
Individual, scenario training 0.41 0.40 0.29
Individual, probability training 0.40 0.36 0.29
Crowd-belief, no training 0.42 0.39 0.30
Crowd-belief, probability training 0.36 0.34 0.23
Team, no training 0.42 0.33 0.22
Team, scenario training 0.36 0.33 0.24
Team, probability training 0.35 0.30 0.19
Year 2
Individual, no training 0.46 0.39 0.26
Individual, probability training 0.42 0.36 0.24
Team, no training 0.38 0.32 0.16
Team, probability training 0.40 0.28 0.16
Superforecasters 0.25 0.19 0.07

The paper reports that both training and teaming were beneficial across all three periods. That is true and it is also the least interesting thing in the table. Convert every cell into a percentage improvement against the untrained individual in the same period and a pattern appears that the paper never states numerically.

Original finding 1: the best lever changes with timing

CEOtudent editorial framework: the timing flip. Percentage reduction in Brier score against the untrained individual in the same period, computed from the Year 1 rows above.

Period Probability training alone Team discussion alone Both together Discussion divided by training
First week 9.1% better 4.5% better 20.5% better 0.50x
Middle 2 weeks 10.0% better 17.5% better 25.0% better 1.75x
Last week 6.5% better 29.0% better 38.7% better 4.50x

In the first week of a question, training was twice the lever that discussion was. By the final week the relationship had inverted almost ninefold in relative terms: discussion was worth 4.5 times what training was worth.

The mechanism is intuitive once the numbers are in front of you. Early on, nobody has information, so the only available edge is reasoning discipline: reference classes, base rates, averaging your own estimates. Late on, information exists in the world and the edge is access to it, which is what a group provides. The paper notes that greater accuracy in teams came from members who gathered and shared information, and that the number of comments an individual posted correlated with their accuracy at r = -0.19 in Year 1 and -0.22 in Year 2, where negative means better.

The practical instruction for an individual is therefore not “use technique X.” It is: when a question is fresh, work on your reasoning; when it is ripening, work on your inputs. Most people do the opposite, brainstorming hardest at the start and going quiet as the deadline nears.

Original finding 2: scenario training is the weakest lever measured

Scenario training in this tournament taught forecasters to generate new futures, entertain more possibilities, use decision trees, and avoid overpredicting change. That is a recognisable description of what most corporate futures work does. It was randomly assigned and separately scored, which makes this the cleanest published head-to-head between scenario-based and probability-based foresight training.

CEOtudent editorial framework: scenario training against probability training. Percentage reduction in Brier score against the no-training condition of the same forecaster type, Year 1.

Period Individuals: scenario Individuals: probability Inside teams: scenario Inside teams: probability
First week 6.8% better 9.1% better 14.3% better 16.7% better
Middle 2 weeks 0.0% (no change) 10.0% better 0.0% (no change) 9.1% better
Last week 6.5% better 6.5% better 9.1% worse 13.6% better

Scenario training lost to probability training in five of the six comparisons and tied in the sixth. In the middle period it produced no measurable improvement at all, for individuals or inside teams. In the final week inside teams it was actively worse than giving the team no training whatsoever, moving the Brier score from 0.22 to 0.24.

The original paper reports the significance test behind the individual comparison: probability training was more effective than scenario training, t(1053) = 2.23, p = .026, and scenario training in turn beat no training, t(1056) = 3.25, p < .001. So scenario training is not worthless. It is simply the weaker of the two, and its advantage disappears in exactly the conditions where most professionals would use it, namely in a group, close to a decision.

One caveat the honest version of this argument has to carry: accuracy is not the only thing scenario planning claims to deliver. Practitioners argue it improves preparedness and organisational conversation rather than point forecasts. That may be true, but it is a different claim, and it is not the one measured here. If you adopt scenario work, adopt it for the thing it was measured on, which is not forecast accuracy.

Original finding 3: selection beats every technique

Look again at the Year 2 row for superforecasters, the top 2% of Year 1 performers placed into elite teams. Against the untrained individual in the same period, they scored 45.7% better in the first week, 51.3% better in the middle, and 73.1% better in the last week. No training or teaming condition anywhere in the table comes close.

This is not an isolated result. Atanasov and colleagues, in work published in the International Journal of Forecasting, compared small elite crowds against larger sub-elite crowds across two crowdsourcing systems over 136 questions.

Verified data. Aggregate Brier scores by crowd type and prediction system (Atanasov et al., International Journal of Forecasting, 2024).

System Sub-elite crowd Elite crowd Elite advantage
Prediction markets 0.215 0.173 20%
Prediction polls 0.203 0.156 23%

Recomputing those percentages from the underlying means gives 19.5% and 23.2%, consistent with the paper’s rounded figures. In a mixed-effects model the authors put the elite effect at b = -0.045, an accuracy improvement of 21%, while the choice between a market and a poll was not significant, b = -0.015, p = .13.

Two conclusions follow, and they are unwelcome in different ways. For an organisation, the highest-return foresight investment is not a methodology purchase; it is finding the handful of people with a track record and asking them. For an individual, there is no shortcut into that group other than the slow one: forecast, record, score, repeat, and find out whether you are actually any good. Note also the superforecasters’ behaviour in the tournament: they made an average of 7.8 predictions per question, against 1.4 to 1.6 for everyone else. Whatever else selection captured, it captured people who update.

The uncomfortable part: the headline evidence is contested

In January 2025, Hauenstein, Thomas, Illingworth and Dougherty published a re-analysis of exactly this data in Psychological Science. Their argument is methodological: forecasters chose which questions to answer and when to answer them, and those choices are not random. Someone in a team may answer different questions, later, than someone forecasting alone, and that difference alone can produce an apparent accuracy advantage that has nothing to do with the discussion.

They modelled latent forecasting ability with item response theory and included these extraneous variables. The effect sizes moved sharply.

CEOtudent editorial framework: what happens to the effects when method variance is controlled. Latent forecast ability effect sizes, Cohen’s d, from Hauenstein et al. (Psychological Science, 2025, Table 4). M2 is the baseline model; M7 is the full method-variance model.

Data set Factor Baseline d (M2) Method-variance d (M7) Change
Year 1 Training +0.050 -0.144 sign reversal
Year 1 Teaming +0.766 -0.285 sign reversal, a shift of 1.051
Year 2, no superforecasters Training +0.174 +0.038 78.2% smaller
Year 2, no superforecasters Teaming +0.256 +0.065 74.6% smaller
Year 2, with superforecasters Training +0.328 +0.227 30.8% smaller
Year 2, with superforecasters Teaming +0.390 +0.247 36.7% smaller

The Year 1 teaming effect is the one to sit with. In the baseline model it is d = +0.766, a large effect and the empirical backbone of a decade of advice about forecasting in groups. Under the full model it is d = -0.285, pointing the other way, a total movement of 1.051 standard deviations.

The authors state their conclusion plainly for the Year 2 non-superforecaster data: asked whether teaming and training affect forecast accuracy, the answer is “no” once the models incorporate item selection and response timing. For Year 2 they report that the best-fitting models do not even include parameters allowing forecasting skill to vary by teaming and training.

Two things are worth keeping straight here, because it is easy to overclaim in either direction. The raw Brier differences in Table 1 are real; nobody disputes that trained team forecasters scored better. What is contested is the causal attribution, whether the training and the teaming produced the improvement or merely accompanied it. And notably, the re-analysis did not dissolve the superforecaster gap; it complicated the interpretation of it, showing that traits associated with strategic responding also discriminate superforecasters from everyone else.

For an individual deciding where to spend effort, this is not a reason for paralysis. It is a reason to weight the methods that hold up under the most scrutiny, which is what the scorecard below does.

Two methods from outside the tournament

Structured group process (Delphi). Rowe and Wright’s review of the evaluative literature counts studies rather than pooling effects. They report that Delphi groups outperformed simple statistical aggregation of the same individuals by 12 studies to 2 with 2 ties, and outperformed traditional unstructured group meetings by 5 to 1. A minor discrepancy worth flagging for anyone citing this: the chapter’s body text says traditional-group comparisons ran “five studies to one, with two ties,” while its own summary says “five to one with one tie.” Vote counting is also a weaker form of evidence than effect-size pooling, so treat this as directional support rather than a measured magnitude. What it supports is specific: structure beats an unstructured meeting, and independent estimates followed by feedback beat a single round.

Reference class forecasting. Flyvbjerg’s work on infrastructure projects documents the baseline problem in unusually hard numbers: average inaccuracy in cost forecasts of 44.7% for rail, 33.8% for bridges and tunnels, and 20.4% for roads, with rail passenger demand forecasts off by an average of -51.4% and 84% of rail projects wrong by more than 20% in either direction. The remedy is to forecast from the distribution of comparable completed projects rather than from the inside of the current one. The UK Department for Transport implementation converted that distribution into required uplifts: for roads, 15% if you will accept a 50% chance of overrun and 45% if you will accept only 10%; for rail, 40% and 68% respectively. The American Planning Association endorsed the method in 2005.

The reason this belongs in a ranking for individuals is that it is the one technique here that a single person can execute in an afternoon with no group, no training programme and no track record, and it attacks the largest documented error in the whole literature.

The scorecard

CEOtudent editorial framework: foresight methods scored for individual use. Evidence strength and robustness are judgements about the published record cited in this article, not new measurements. Solo-runnable and cost are practical assessments.

Method Best measured effect Survives re-analysis Runnable alone Cost Verdict
Track-record selection 73.1% error reduction (tournament); 20 to 23% (elite crowds) Gap survives; interpretation contested Only by becoming the record High, measured in years Do it anyway. Nothing else is close.
Reference class forecasting Attacks documented 20 to 51% baseline errors Not challenged in the literature reviewed Yes, fully Low, one afternoon per decision Start here. Best return per hour.
Probability training 9.1 to 10.0% early in a question Substantially reduced or reversed under method-variance modelling Yes Low, one module took about 45 minutes Worth it for the habits, not the headline number.
Structured group process Delphi beat statistical aggregation 12 to 2 with 2 ties Vote counting only, no pooled effect sizes No, needs three or more people Medium Use when you have the people.
Team discussion 29.0% late in a question Reduced to near zero, sign reversed in Year 1 No Medium Use for information access, not as a technique.
Prediction markets Statistically tied with team polls Not separately challenged No High, needs a crowd No individual case.
Crowd-belief exposure 2.5 to 4.5% from seeing others’ numbers alone Not separately modelled Yes, if a public forecast exists Very low Free, small, take it.
Scenario planning and training 0.0% in the middle period; 9.1% worse inside teams late Not the strongest claim it makes Yes Medium to high Not for accuracy. Adopt only for what it was measured on.

What to actually do

A defensible personal foresight practice from this evidence has four parts and none of them require a workshop.

  1. Build a reference class before you build a forecast. Whatever you are predicting, find the ten to thirty most comparable completed cases and look at what actually happened to them. This attacks the biggest measured error in the literature and needs nobody’s permission.
  2. Split your effort by stage. Early in any open question, spend your time on reasoning discipline: base rates, multiple independent estimates, averaging. Later, spend it on gathering and exchanging information. The measured leverage ratio between these two moves from 0.50x to 4.50x across the life of a question.
  3. Keep score, in writing, with numbers. This is the only route into the group with the largest measured advantage. A forecast without a recorded probability and a resolution date cannot be scored, and an unscored forecaster cannot improve. Update often; the superforecasters updated roughly five times more than anyone else.
  4. Demote scenario exercises to what they measured well on. If you use them, use them to widen the set of possibilities you have considered, not to sharpen a probability. On sharpening a probability, they measured worst here.

Lead the practice like a CEO: pick the two methods with the best measured return and refuse the rest, however fashionable. Learn like a student: keep the scorecard honest even when it says your favourite technique did nothing, because the whole value of this literature is that it was willing to publish that result about itself. Twice.

For the upstream question of noticing what to forecast in the first place, see how to spot weak signals; for how a decade of confident expert predictions actually resolved, see the forecast scorecard.

Frequently asked questions

Does the 2025 re-analysis mean training and teams are useless?
No, and the paper does not claim that. It claims the original causal conclusions do not hold once question selection and response timing are modelled. Trained team forecasters still recorded better raw scores. What is in doubt is whether the training and the teaming caused it. Practically, that argues for weighting methods that were not built on that particular inference, which is why reference class forecasting ranks above probability training here.

Why is scenario planning so popular if it measured worst?
Because it is evaluated on a different outcome. Scenario work is bought for preparedness, alignment and the quality of the internal conversation, and it may well deliver those. This article measures forecast accuracy only, and on that outcome it was the weakest lever tested.

Is superforecasting just talent?
The tournament evidence points more at behaviour than at innate ability: superforecasters made an average of 7.8 predictions per question versus 1.4 to 1.6 for other conditions, and the identification itself was a track record, not a test. The 2025 re-analysis complicates this by showing that response-strategy traits also separate the group, so the honest answer is that scoring identifies something real whose composition is not settled.

How many comparable cases do I need for a reference class?
The literature does not fix a number, and any figure we gave you would be invented. What the UK Department for Transport implementation shows is the shape of the requirement: enough completed, genuinely comparable cases to see a distribution rather than a point, established statistically as similar in risk before use.

Can I run a Delphi process with two colleagues?
The core mechanism, independent estimates collected before anyone hears anyone else, then feedback, then revision, works at small scale and costs almost nothing. The evidence base compares panels rather than pairs, so treat a three-person version as a sensible application of the principle rather than as something measured.

Sources

  • Mellers, Ungar, Baron, Ramos, Gurcay, Fincher, Scott, Moore, Atanasov, Swift, Murray, Stone and Tetlock, Psychological Strategies for Winning a Geopolitical Forecasting Tournament, Psychological Science, 2014, volume 25, issue 5, pages 1106 to 1115. Table 1 Brier scores by condition and period; the probability versus scenario training significance tests; comment count correlations; prediction counts per question.
  • Hauenstein, Thomas, Illingworth and Dougherty, Rethinking the Role of Teams and Training in Geopolitical Forecasting: The Effect of Uncontrolled Method Variance on Statistical Conclusions, Psychological Science, 2025, volume 36, issue 1, pages 3 to 18. Item response theory re-analysis; Table 4 latent ability effect sizes for the baseline and full method-variance models; the conclusion for Year 2 non-superforecaster data.
  • Atanasov, Witkowski, Mellers and Tetlock, Crowd Prediction Systems: Markets, Polls, and Elite Forecasters, International Journal of Forecasting, 2024. Aggregate Brier scores by crowd type and system across 136 questions; the mixed-effects elite coefficient and the non-significant market versus poll comparison.
  • Rowe and Wright, Expert Opinions in Forecasting: The Role of the Delphi Technique, in Principles of Forecasting, Kluwer Academic Publishers, 2001. Vote counts of Delphi against statistical and traditional groups, and the principles for running a Delphi panel.
  • Flyvbjerg, From Nobel Prize to Project Management: Getting Risks Right, Project Management Journal, 2006. Average cost and demand forecast inaccuracy by project type; the UK Department for Transport optimism bias uplifts at the 50% and 80% levels; the 2005 American Planning Association endorsement of reference class forecasting.

This content was compiled with the support of AI following in-depth research, then written and prepared for publication by the CEOtudent editorial team.

This post is also available in: Türkçe Français Español Deutsch

Benzer içerikler