GelişimStrateji
0

Intellectual Honesty as a Competitive Advantage: How to Be Right More Often by Being Wrong Faster

Professional smiling while revising a note in a paper notebook at a sunlit desk

TL;DR. Intellectual honesty sounds like a character trait. In forecasting research it behaves like a skill with a score. In the Good Judgment Project, a research team that competed in a US intelligence-funded forecasting tournament, the most accurate forecasters “made frequent, small updates”, while low-skill forecasters “were prone to confirm initial judgments or make infrequent, large revisions” (Atanasov and colleagues, 2020, drawing on more than 400,000 predictions on almost 500 questions). The elite group, the superforecasters, submitted 7.8 forecasts per question on average, against 1.6 for regular teams and 1.4 for people forecasting alone. We recalculated the published accuracy table: in the final week before questions closed, superforecasters’ Brier score of 0.07 was 56% lower than that of trained teams (0.16) and 73% lower than that of untrained individuals (0.26). Two honest caveats come with this. Updating wins only when the evidence is informative: in one study, the most open-minded people did worse on upsets. And a 2024 reanalysis argues that the celebrated training and teaming effects shrink or disappear once response timing and question choice are controlled. The practical rule survives both caveats: commit to a number, state in advance what would change it, and move it in small steps as evidence arrives.

Why intellectual honesty is a performance variable, not a virtue

Intellectual honesty is usually framed as ethics: admit mistakes because it is right. That framing makes it optional, and busy, ambitious people tend to treat admitting error as a cost to their credibility.

The forecasting literature reframes the question. If you express beliefs as probabilities and score them against what actually happens, then the habits we call intellectual honesty (noticing disconfirming evidence, revising in public, not clinging to a first answer) leave a measurable trace in accuracy. They stop being a virtue you display and become an input you can track.

Three bodies of evidence make the case, each with a limit worth knowing:

  • Belief updating in real-world forecasting (Mellers and colleagues, 2014 and 2015; Atanasov and colleagues, 2020): how often and how much the best forecasters revise.
  • Actively open-minded thinking (Haran, Ritov and Mellers, 2013): the disposition to seek out information, including information that could prove you wrong, and what it does to accuracy.
  • Intellectual humility (Leary and colleagues, 2017): the degree to which people recognise that their beliefs might be wrong, and how that shapes the way they judge arguments and other people.

This is the CEOtudent lens in one sentence. The CEO owns a decision and commits to a position. The student treats that position as a draft and keeps revising it. Accuracy comes from doing both at once.

How accuracy gets measured: the Brier score in plain terms

Forecasting research needs a scoring rule that rewards honesty, and the standard one is the Brier score, proposed by Brier in 1950 in the journal Monthly Weather Review and used throughout the Good Judgment Project. In the version the project used, a forecast on a yes-or-no question is scored as the sum of squared gaps between the probabilities you gave and what happened. A perfect score is 0 and the worst possible is 2.

Mellers and colleagues give the worked example. Say 90% that an event will happen. If it happens, your score is (0.9 – 1)² + (0.1 – 0)² = 0.02. If it does not, your score is (0.9 – 0)² + (0.1 – 1)² = 1.62. Being confidently wrong costs 81 times as much as being confidently right earns. Scores are averaged across every day a question stays open, so a forecast left stale for weeks keeps adding to your error.

Two properties matter for intellectual honesty. First, the rule is “proper”: the best long-run strategy is to report what you actually believe, not to exaggerate your confidence. Second, because the score is averaged over days, updating is visible. A forecaster who is right in the end but refused to move for most of the question’s life still pays for every stale day.

What the best forecasters actually do differently

The Good Judgment Project ran randomised experiments inside the tournament. In Year 2 the top 2% of Year 1 performers were placed in elite teams of superforecasters. The published table of accuracy by time period is the most useful starting point, because it shows accuracy at the start, the middle and the end of each question.

Table 1. Average Brier scores by time period, Year 2 (verified data from Mellers and colleagues, 2014, Table 1; lower is better; questions open at least a month)

Condition First week Middle 2 weeks Last week
Individual forecasters, no training 0.46 0.39 0.26
Individual forecasters, probability training 0.42 0.36 0.24
Team forecasters, no training 0.38 0.32 0.16
Team forecasters, probability training 0.40 0.28 0.16
Superforecasters 0.25 0.19 0.07

The same paper reports how much each group engaged. Across Year 2, participants made an average of 1.8 predictions per question. Independent forecasters made 1.4, regular teams 1.6, and superforecasters 7.8. The authors call the superforecasters’ engagement “extraordinary”. Superforecasters also answered more questions: 95 on average, against 61 for regular teams.

The paper also offers a concrete example of what failing to update costs. Calibration was worst for forecasts of 100%, and the authors put the cause plainly: “The problem was a lack of updating.” Forecasts of 100% made in the first 20% of a question’s life came true only about 70% of the time. The same forecasts made in the last 20% of days were right about 90% of the time. Certainty stated early behaved like a guess held too long.

The follow-up paper on superforecasters (Mellers and colleagues, 2015) found that they maintained high accuracy for two years running, “defying expectations of regression toward the mean”. The authors support four mutually reinforcing explanations: cognitive abilities and styles, task-specific skills, motivation and commitment, and enriched environments. Their conclusion is that superforecasters are “partly discovered and partly created”. That second half is the reason this matters to anyone who is not already a superforecaster.

Small steps beat big jumps: the updating evidence

The most direct evidence on updating comes from Atanasov, Witkowski, Ungar, Mellers and Tetlock (2020). They used four years of tournament data and separated three aspects of updating: frequency (how often a forecaster revises), magnitude (the average size of each revision) and confirmation propensity (the tendency to resubmit the same forecast).

Their findings, in the authors’ own summary:

  • The most accurate forecasters made frequent, small updates.
  • Low-skill forecasters were prone to confirm initial judgments or make infrequent, large revisions.
  • High-frequency updaters scored higher on crystallized intelligence and open-mindedness, accessed more information, and improved over time.
  • Small-increment updaters had higher fluid intelligence scores and derived their advantage from their initial forecasts.
  • Update magnitude mediated the causal effect of training on accuracy.
  • Frequent, small revisions provided “reliable and valid signals of skill”.

Two points deserve emphasis. First, frequency and magnitude are different habits with different sources. Updating often reflects engagement and information gathering. Updating in small steps reflects starting from a reasonable estimate so that no large correction is needed. Second, the authors frame both as signals organisations can use to identify people who manage uncertainty well. In other words, the way someone changes their mind is itself evidence of their judgment.

A calculation: what anchoring, overconfidence and incremental updating cost

To make the mechanism concrete, we built a transparent illustration. It uses the same two-outcome Brier score as the Good Judgment Project, averaged over the days a question is open. The question runs for 10 days. The evidence that arrives over those days points increasingly towards “yes”. Four forecasters respond differently. All numbers in Table 2 are illustrative: they show how the scoring rule treats each behaviour, not how any real forecaster performed.

Table 2. Four updating styles on an illustrative 10-day question (CEOtudent calculation; illustrative numbers; two-outcome Brier score averaged over days, 0 best, 2 worst)

Style Forecast path Updates Mean update size Score if “yes” happens Score if upset (“no”) Expected score*
Anchored 30% for all 10 days 0 0 points 0.980 0.180 0.820
Overconfident 95% from day 1, never revised 0 0 points 0.005 1.805 0.545
Late big jump 30% for 7 days, then 90% 1 60 points 0.692 0.612 0.676
Incremental 30, 35, 42, 50, 57, 63, 70, 78, 85, 90% 9 6.7 points 0.398 0.798 0.478

*Expected score assumes the evidence trend turns out to be right 80% of the time. For the overconfident forecaster, whose view is set before the evidence arrives, we assume the first impression is right 70% of the time, which echoes the Mellers and colleagues finding that early 100% forecasts came true about 70% of the time. Both percentages are illustrative assumptions.

What the calculation shows:

  1. Anchoring is the most expensive habit when things change. The anchored forecaster is punished every day the world moves away from their starting view.
  2. Overconfidence wins big or loses catastrophically. A score of 0.005 when right, 1.805 when wrong. That is the asymmetry of the Brier score in action.
  3. A late big jump is better than no jump, but pays for every stale day before it. This is the “infrequent, large revisions” pattern Atanasov and colleagues associate with lower skill.
  4. Incremental updating has the best expected score, but it is not free. If the evidence misleads, the incremental updater does worse than the anchored one (0.798 against 0.180). Updating is a bet that the evidence is informative.

That fourth point is not a quirk of our example. It matches the most careful study of open-mindedness and accuracy, discussed next.

The honest limits: when open-mindedness backfires, and the reanalysis

Haran, Ritov and Mellers (2013) tested four thinking styles as predictors of accuracy across three studies: actively open-minded thinking, need for cognition, grit, and the tendency to maximise. Only actively open-minded thinking predicted performance. The mechanism was information: more open-minded people collected more information before estimating, and that extra information produced the accuracy gain. When the amount of available information was held constant (Study 2), open-mindedness had no effect on performance.

The third study, which asked people to predict American football results, adds the caveat. When outcomes matched the pre-game information, open-minded thinkers were more accurate. When the result was an upset, higher open-mindedness went with worse performance, because those forecasters had followed the information more closely. As the authors put it, when information was misleading, highly open-minded people were “more susceptible to invalid information”. Their conclusion is conditional: as long as information is at least somewhat predictive, open-minded thinking should help.

The Good Judgment Project paper has its own example. On a question about a lethal confrontation in the South or East China Sea, the best forecasters started near 20%, the base rate, and lowered their estimates over time. Then a South Korean coast guard officer was killed by a Chinese fisherman. Trained teams scored 1.44 on that question, worse than untrained independent forecasters at 1.31.

The larger caveat is methodological. In a reanalysis published in Psychological Science, Hauenstein, Thomas, Illingworth and Dougherty (2024) modelled the same tournament data with item response theory and controlled for item difficulty, the timing of forecasts and which questions forecasters chose to answer. In their words, the best-fitting models “substantially eliminated, reduced, and, in some cases, even reversed the effects” of the teaming and training manipulations on latent forecasting ability. They also report that these extraneous variables can discriminate between superforecasters and others, which makes it harder to say what underlies superforecasters’ performance.

How this affects the argument here: the reanalysis challenges whether a 45-minute training module or team membership causes better forecasting. It does not overturn the descriptive finding that the most accurate forecasters update frequently and in small steps. Response timing, one of the variables the reanalysis controls for, is itself partly an updating behaviour. Mellers and colleagues had reported in 2014 that training and teaming remained significant after controlling for the timing and number of forecasts. The two teams disagree on method, and a careful reader should hold the causal claims about training loosely.

Intellectual humility and how you judge people who change their minds

Intellectual honesty also has a social side. Leary and colleagues (2017) developed an Intellectual Humility Scale, defining the trait as the degree to which people recognise that their beliefs might be wrong. In the first of four studies, intellectual humility was associated with openness, curiosity, tolerance of ambiguity and low dogmatism. People high in intellectual humility were less inclined to see politicians who changed their attitudes as “flip-flopping”, and they were more attuned to the strength of persuasive arguments than people low in the trait.

This matters for leaders. If the culture you set treats every revision as flip-flopping, you are training your team towards the anchored and late-big-jump styles in Table 2. If revisions in response to evidence are treated as normal, you make small, frequent updates socially affordable. The forecasting evidence says that is where accuracy lives. For a way to tell whether your own first answer is sticking for the wrong reasons, see our piece on first-conclusion bias and why people accept an AI’s first answer.

The Intellectual Honesty Operating Protocol

Evidence and framework need to stay separate. The findings above are evidence. The protocol below is our editorial synthesis: a way to turn those findings into weekly habits. Each step is mapped to the finding that motivates it, and to the limit of that finding.

Table 3. The Intellectual Honesty Operating Protocol (CEOtudent editorial framework, mapped to verified findings)

Step What you do Finding that motivates it Limit to respect
1. Commit to a number State important beliefs as probabilities (“65% this launch hits target”), not adjectives The Brier score rewards honest probabilities and makes accuracy trackable (Mellers and colleagues, 2014) A number without a later check is theatre; step 5 is required
2. Pre-commit to what would change your mind Before deciding, write the two or three observations that would move your number by 10 points or more Early certainty was right only about 70% of the time; “the problem was a lack of updating” (Mellers and colleagues, 2014) Pick signals that are genuinely predictive, not just visible
3. Start from a base rate Anchor the first estimate on how often similar things happen, then adjust Small-increment updaters drew their advantage from better initial forecasts (Atanasov and colleagues, 2020) Base rates can mislead if your case is unusual
4. Update often, in small steps Revisit live beliefs on a schedule; move in steps sized to the evidence, typically a few points The most accurate forecasters made frequent, small updates; low-skill ones confirmed or jumped (Atanasov and colleagues, 2020) Updating is a bet that evidence is informative (Haran and colleagues, 2013)
5. Keep an error log and score it Record forecasts and outcomes; compute a simple Brier score each quarter Frequent, small revisions were “reliable and valid signals of skill” (Atanasov and colleagues, 2020) Few decisions per year means noisy scores; judge trends, not single results
6. Seek information that could prove you wrong Before a big decision, spend deliberate time on the strongest counter-evidence Open-minded thinkers collected more information, which explained their accuracy (Haran and colleagues, 2013) Extra information helps only if it is valid
7. Make revision socially safe Praise evidence-based changes of mind in your team; never punish them as flip-flopping High intellectual humility goes with less “flip-flopping” judgment and more attention to argument strength (Leary and colleagues, 2017) Correlational evidence; culture effects are our inference

Several tools on CEOtudent support individual steps. For step 1, see how to think in bets. For step 2, the kill criteria guide shows how to set conditions in advance. For step 5, the decision journal protocol gives you a template for logging forecasts and reviewing them. For step 6, a pre-mortem makes counter-evidence a scheduled part of planning.

The CEO and the student

The career argument for intellectual honesty is simple once accuracy is measurable. People who manage uncertainty well are valuable, and Atanasov and colleagues suggest that the pattern of someone’s updates is a reliable signal of that skill. An organisation that tracks forecasts can see who revises in small, timely steps and who defends a first answer until reality forces a jump.

The CEO side of the lens is ownership. You commit to a number, put your name on it, and decide. Intellectual honesty does not mean avoiding commitment. The anchored forecaster in Table 2 committed, and so did the overconfident one. The difference is what happens after the commitment.

The student side is revision. You treat every position as the current draft, you know in advance what would change it, and you change it by the right amount when that evidence arrives. The data suggest this is not a trade-off between confidence and humility. The most accurate forecasters in the Good Judgment Project did both: they committed to numbers and kept moving them.

Being wrong faster does not mean being wrong more often. It means finding out sooner, while the correction is small and cheap.

FAQ

What is intellectual honesty in decision-making?
In practical terms, it is the habit of stating beliefs precisely enough to be proven wrong, looking for evidence that could prove them wrong, and revising them in proportion to that evidence. Forecasting research makes this measurable by scoring probability judgments against outcomes.

Do people who change their minds often make better predictions?
In the Good Judgment Project data, yes, when the changes are frequent and small. Atanasov and colleagues (2020) found that the most accurate forecasters made frequent, small updates, while low-skill forecasters confirmed their first judgments or made infrequent, large revisions. Superforecasters submitted 7.8 forecasts per question on average, against 1.6 for regular teams.

Is being open-minded always better?
No. Haran, Ritov and Mellers (2013) found that actively open-minded thinkers were more accurate because they gathered more information, but on upsets, where the information pointed the wrong way, higher open-mindedness went with worse performance. Updating helps when the evidence is at least somewhat predictive.

Can forecasting skill be trained?
The original Good Judgment Project results said a roughly 45-minute probability training module improved accuracy, with benefits lasting across forecasting seasons of about 8 to 10 months each. A 2024 reanalysis by Hauenstein and colleagues argues that these training and teaming effects shrink, disappear or reverse once timing, question choice and question difficulty are controlled. The causal question is contested. The link between updating style and accuracy is better supported.

How can I measure my own accuracy?
Write down forecasts as probabilities with a date and a resolution rule, then score them. For a yes-or-no question in the two-outcome version used here, the score is the squared gap on “yes” plus the squared gap on “no”. Saying 90% on something that happens scores 0.02, and saying 90% on something that does not scores 1.62. Average the scores over a quarter and track the trend.

Sources

  • Mellers, B., Ungar, L., Baron, J., Ramos, J., Gurcay, B., Fincher, K., Scott, S. E., Moore, D., Atanasov, P., Swift, S. A., Murray, T., Stone, E. and Tetlock, P. E. (2014). Psychological Strategies for Winning a Geopolitical Forecasting Tournament. Psychological Science, 25(5). Table 1 (Brier scores by time period), engagement measures (predictions per question), calibration discussion, Brier score worked example, South and East China Sea example.
  • Mellers, B., Stone, E., Murray, T., Minster, A., Rohrbaugh, N., Bishop, M., Chen, E., Baker, J., Hou, Y., Horowitz, M., Ungar, L. and Tetlock, P. (2015). Identifying and Cultivating Superforecasters as a Method of Improving Probabilistic Predictions. Perspectives on Psychological Science, 10(3), 267-281. Abstract: persistence of superforecaster accuracy and four explanations.
  • Atanasov, P., Witkowski, J., Ungar, L., Mellers, B. and Tetlock, P. (2020). Small steps to accuracy: Incremental belief updaters are better forecasters. Organizational Behavior and Human Decision Processes, 160, 19-35. Abstract findings on update frequency, magnitude and confirmation; conference version at the 21st ACM Conference on Economics and Computation (EC‘20) for dataset description and update measures.
  • Haran, U., Ritov, I. and Mellers, B. A. (2013). The role of actively open-minded thinking in information acquisition, accuracy, and calibration. Judgment and Decision Making, 8(3), 188-201. Studies 1 to 3, mediation by information acquisition, upset results, general discussion.
  • Hauenstein, C., Thomas, R., Illingworth, D. and Dougherty, M. (2024, published online; 2025 print issue). Rethinking the Role of Teams and Training in Geopolitical Forecasting: The Effect of Uncontrolled Method Variance on Statistical Conclusions. Psychological Science, 36(1), 3-18. Author accepted manuscript: abstract and discussion.
  • Leary, M. R., Diebels, K. J., Davisson, E. K., Jongman-Sereno, K. P., Isherwood, J. C., Raimi, K. T., Deffler, S. A. and Hoyle, R. H. (2017). Cognitive and Interpersonal Features of Intellectual Humility. Personality and Social Psychology Bulletin, 43(6), 793-813. Abstract: four studies using the Intellectual Humility Scale.

Table 1 reproduces published data from Mellers and colleagues (2014). The percentage comparisons in the text (for example “56% lower”) are CEOtudent calculations from that table. Table 2 is a CEOtudent calculation with illustrative numbers. Table 3 is a CEOtudent editorial framework.


This content was compiled with the support of AI following in-depth research, then written and prepared for publication by the CEOtudent editorial team.

This post is also available in: Türkçe Français Español Deutsch

Benzer içerikler