<\/span><\/h2>\nBrier scores measure the accuracy of probabilistic forecasts. Lower is better. A forecaster who says 90% and is right scores 0.02 on the scale used in the paper; one who says 90% and is wrong scores 1.62.<\/p>\n
Verified data. Average Brier score by condition and by period in a question’s life (Mellers et al., Psychological Science, 2014, Table 1). Only questions open at least a month are included.<\/strong><\/p>\n\n\n\n| Condition<\/th>\n | First week<\/th>\n | Middle 2 weeks<\/th>\n | Last week<\/th>\n<\/tr>\n<\/thead>\n |
\n\nYear 1<\/strong><\/td>\n| <\/td>\n | <\/td>\n | <\/td>\n<\/tr>\n | \n| Individual, no training<\/td>\n | 0.44<\/td>\n | 0.40<\/td>\n | 0.31<\/td>\n<\/tr>\n | \n| Individual, scenario training<\/td>\n | 0.41<\/td>\n | 0.40<\/td>\n | 0.29<\/td>\n<\/tr>\n | \n| Individual, probability training<\/td>\n | 0.40<\/td>\n | 0.36<\/td>\n | 0.29<\/td>\n<\/tr>\n | \n| Crowd-belief, no training<\/td>\n | 0.42<\/td>\n | 0.39<\/td>\n | 0.30<\/td>\n<\/tr>\n | \n| Crowd-belief, probability training<\/td>\n | 0.36<\/td>\n | 0.34<\/td>\n | 0.23<\/td>\n<\/tr>\n | \n| Team, no training<\/td>\n | 0.42<\/td>\n | 0.33<\/td>\n | 0.22<\/td>\n<\/tr>\n | \n| Team, scenario training<\/td>\n | 0.36<\/td>\n | 0.33<\/td>\n | 0.24<\/td>\n<\/tr>\n | \n| Team, probability training<\/td>\n | 0.35<\/td>\n | 0.30<\/td>\n | 0.19<\/td>\n<\/tr>\n | \nYear 2<\/strong><\/td>\n| <\/td>\n | <\/td>\n | <\/td>\n<\/tr>\n | \n| Individual, no training<\/td>\n | 0.46<\/td>\n | 0.39<\/td>\n | 0.26<\/td>\n<\/tr>\n | \n| Individual, probability training<\/td>\n | 0.42<\/td>\n | 0.36<\/td>\n | 0.24<\/td>\n<\/tr>\n | \n| Team, no training<\/td>\n | 0.38<\/td>\n | 0.32<\/td>\n | 0.16<\/td>\n<\/tr>\n | \n| Team, probability training<\/td>\n | 0.40<\/td>\n | 0.28<\/td>\n | 0.16<\/td>\n<\/tr>\n | \n| Superforecasters<\/td>\n | 0.25<\/td>\n | 0.19<\/td>\n | 0.07<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n The paper reports that both training and teaming were beneficial across all three periods. That is true and it is also the least interesting thing in the table. Convert every cell into a percentage improvement against the untrained individual in the same period and a pattern appears that the paper never states numerically.<\/p>\n <\/span>Original finding 1: the best lever changes with timing<\/span><\/h2>\nCEOtudent editorial framework: the timing flip. Percentage reduction in Brier score against the untrained individual in the same period, computed from the Year 1 rows above.<\/strong><\/p>\n\n\n\n| Period<\/th>\n | Probability training alone<\/th>\n | Team discussion alone<\/th>\n | Both together<\/th>\n | Discussion divided by training<\/th>\n<\/tr>\n<\/thead>\n | \n\n| First week<\/td>\n | 9.1% better<\/td>\n | 4.5% better<\/td>\n | 20.5% better<\/td>\n | 0.50x<\/td>\n<\/tr>\n | \n| Middle 2 weeks<\/td>\n | 10.0% better<\/td>\n | 17.5% better<\/td>\n | 25.0% better<\/td>\n | 1.75x<\/td>\n<\/tr>\n | \n| Last week<\/td>\n | 6.5% better<\/td>\n | 29.0% better<\/td>\n | 38.7% better<\/td>\n | 4.50x<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n In the first week of a question, training was twice the lever that discussion was. By the final week the relationship had inverted almost ninefold in relative terms: discussion was worth 4.5 times what training was worth.<\/p>\n The mechanism is intuitive once the numbers are in front of you. Early on, nobody has information, so the only available edge is reasoning discipline: reference classes, base rates, averaging your own estimates. Late on, information exists in the world and the edge is access to it, which is what a group provides. The paper notes that greater accuracy in teams came from members who gathered and shared information, and that the number of comments an individual posted correlated with their accuracy at r = -0.19 in Year 1 and -0.22 in Year 2, where negative means better.<\/p>\n The practical instruction for an individual is therefore not “use technique X.” It is: when a question is fresh, work on your reasoning; when it is ripening, work on your inputs. Most people do the opposite, brainstorming hardest at the start and going quiet as the deadline nears.<\/p>\n <\/span>Original finding 2: scenario training is the weakest lever measured<\/span><\/h2>\nScenario training in this tournament taught forecasters to generate new futures, entertain more possibilities, use decision trees, and avoid overpredicting change. That is a recognisable description of what most corporate futures work does. It was randomly assigned and separately scored, which makes this the cleanest published head-to-head between scenario-based and probability-based foresight training.<\/p>\n CEOtudent editorial framework: scenario training against probability training. Percentage reduction in Brier score against the no-training condition of the same forecaster type, Year 1.<\/strong><\/p>\n\n\n\n| Period<\/th>\n | Individuals: scenario<\/th>\n | Individuals: probability<\/th>\n | Inside teams: scenario<\/th>\n | Inside teams: probability<\/th>\n<\/tr>\n<\/thead>\n | \n\n| First week<\/td>\n | 6.8% better<\/td>\n | 9.1% better<\/td>\n | 14.3% better<\/td>\n | 16.7% better<\/td>\n<\/tr>\n | \n| Middle 2 weeks<\/td>\n | 0.0% (no change)<\/td>\n | 10.0% better<\/td>\n | 0.0% (no change)<\/td>\n | 9.1% better<\/td>\n<\/tr>\n | \n| Last week<\/td>\n | 6.5% better<\/td>\n | 6.5% better<\/td>\n | 9.1% worse<\/strong><\/td>\n| 13.6% better<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n Scenario training lost to probability training in five of the six comparisons and tied in the sixth. In the middle period it produced no measurable improvement at all, for individuals or inside teams. In the final week inside teams it was actively worse than giving the team no training whatsoever, moving the Brier score from 0.22 to 0.24.<\/p>\n The original paper reports the significance test behind the individual comparison: probability training was more effective than scenario training, t(1053) = 2.23, p = .026, and scenario training in turn beat no training, t(1056) = 3.25, p < .001. So scenario training is not worthless. It is simply the weaker of the two, and its advantage disappears in exactly the conditions where most professionals would use it, namely in a group, close to a decision.<\/p>\n One caveat the honest version of this argument has to carry: accuracy is not the only thing scenario planning claims to deliver. Practitioners argue it improves preparedness and organisational conversation rather than point forecasts. That may be true, but it is a different claim, and it is not the one measured here. If you adopt scenario work, adopt it for the thing it was measured on, which is not forecast accuracy.<\/p>\n | | | | | |