GelişimStrateji
0

How to Develop Good Judgment: What Research on Expert Intuition Actually Shows

Person pausing at a bright window with a notebook and a chessboard, thinking before making a decision

TL;DR: “Trust your gut” and “ignore your gut” are both wrong, because whether intuition works is not a fact about you. It is a fact about the environment you learned in. Two researchers who spent careers disagreeing about intuition, Daniel Kahneman and Gary Klein, sat down together and agreed on the boundary: skilled intuition requires an environment regular enough to contain real cues, plus prolonged practice with feedback good enough to learn them. Where both hold, experts are excellent. Where either fails, experience produces confidence without accuracy, and confidence does not warn you which case you are in. Below: the verified evidence, a scorecard to grade your own domain on four dimensions, and the specific development path for people whose field fails the test, which is most knowledge work. Lead like a CEO who knows which decisions to trust, learn like a student who checks the answer.

Ask a room of accomplished people how they make their hardest calls and most will say some version of the same thing: experience, pattern recognition, a feel for it. Ask them how they know their feel is calibrated and the room gets quiet. That silence is the subject of this piece.

The question matters more now than it did five years ago. As more analysis and execution shifts to machines, what stays on the human side of the line is judgment: knowing which problem to solve, which output to trust, when the confident answer in front of you is wrong. We argued the economic case for this in The Judgment Economy. This piece is the developmental one. If judgment is the asset, how is it actually built, and what does the research say about people who think they have built it but have not?

The agreement that settled the question

For decades, two research traditions fought about intuition. The heuristics-and-biases school, associated with Daniel Kahneman, documented systematic error in expert judgment. The naturalistic decision-making school, associated with Gary Klein, documented fireground commanders and nurses making superb split-second calls that no formal analysis could match.

In 2009 they published the result of an adversarial collaboration in American Psychologist, titled “Conditions for Intuitive Expertise: A Failure to Disagree.” The finding is that both were right about different environments, and the boundary between them is specifiable. Evaluating whether an intuitive judgment can be trusted requires assessing two things: the predictability of the environment, and whether the individual had an opportunity to learn its regularities.

Put plainly, skilled intuition needs two conditions, and needs both:

Condition one: the environment contains valid cues. There must be genuine regularity to detect. Chess positions, weather systems and the physics of a burning building contain real, repeating structure. Next quarter’s market and the trajectory of a particular startup contain much less.

Condition two: you had the chance to learn those cues. Prolonged practice, with feedback of high enough quality and speed to attach outcomes to the decisions that produced them.

Miss condition one and no amount of experience helps, because there is nothing stable to learn. Miss condition two and the regularity exists but you never received the signal that would teach it to you.

What the record actually shows

The following are published findings, reported as their authors reported them.

Verified findings from the expertise and judgment literature

Source What it examined What it established
Kahneman and Klein (2009), American Psychologist, 64(6), 515-526 Boundary conditions for trustworthy intuition Two necessary conditions: an environment of sufficient regularity, and the opportunity to learn its regularities through prolonged practice with high-quality feedback
Shanteau (1992), Organizational Behavior and Human Decision Processes Which professions show accurate expert judgment and which do not Strong performance: weather forecasters, livestock judges, test pilots, soil judges, chess masters, accountants, astronomers, physicists, mathematicians, grain inspectors, photo interpreters, insurance analysts
Shanteau (1992), same study The contrasting list Weak performance, often little better than novices: clinical psychologists, psychiatrists, admissions officers, court judges, behavioural researchers, parole officers, and also noted in stockbroking, personnel selection and intelligence analysis
Macnamara, Hambrick and Oswald (2014), Psychological Science Meta-analysis of how much deliberate practice explains performance Variance explained: games 26%, music 21%, sports 18%, education 4%, professions less than 1%
Hogarth, Lejarraga and Soyer (2015), Current Directions in Psychological Science Kind versus wicked learning environments In kind environments feedback links outcomes accurately to the actions that caused them; in wicked ones feedback is missing, delayed or misleading, and confidence does not distinguish the two
Tetlock, Expert Political Judgment (2005) Long-running study of expert forecasting 284 experts, roughly 28,000 forecasts; “foxes” who draw on many frameworks outperformed “hedgehogs” committed to one, especially over longer horizons
Good Judgment Project, IARPA forecasting tournament (2011-2015) Open competition in geopolitical forecasting The project won the tournament; its best amateur forecasters outperformed professional analysts, demonstrating that forecasting accuracy is trainable and measurable

Read the Shanteau lists side by side and the pattern jumps out. The strong list is full of domains with fast, unambiguous, repeated feedback: the weather either happened or it did not, the aircraft flew or it did not, the position was won or lost. The weak list is full of domains where the outcome arrives years later, is contaminated by a hundred other causes, or never arrives in a form anyone records. Parole decisions and psychiatric prognoses do not come with scoreboards.

Now read the Macnamara row against them. Deliberate practice explained 26 percent of performance variance in games and less than 1 percent in professions. The standard interpretation is that talent matters more than the ten-thousand-hours story suggested. There is a second reading that follows directly from Kahneman and Klein: games are the archetypal high-validity, fast-feedback environment, and “professions” is a bucket dominated by low-validity ones. Practice pays where the environment teaches. This connection is our editorial inference across the two literatures, not a claim either study made, but it is the reading that turns the finding into advice.

The most unsettling result in the whole set is Hogarth’s. Confidence does not track environment quality. People who learned in wicked environments feel exactly as sure as people who learned in kind ones. Your certainty is not evidence about your accuracy, which is why this cannot be settled by introspection and needs a scorecard instead.

The Judgment Validity Scorecard

Below is a diagnostic for grading a specific recurring decision you make, not your whole career. Judgment is decision-specific: the same executive can have superb intuition about which engineer will succeed and terrible intuition about which market will open.

Pick one decision you make repeatedly. Score each dimension 0, 1 or 2.

CEOtudent Judgment Validity Scorecard (editorial framework, built on the Kahneman-Klein conditions and Hogarth’s feedback criteria)

Dimension Score 0 Score 1 Score 2
Regularity. Does the situation repeat with stable structure? Each case is close to unique; the rules change between instances Broad patterns recur but the specifics shift a lot Situations repeat with recognisable structure
Feedback speed. How long until you learn the outcome? Years, or never Months Days or faster
Feedback clarity. Can you attribute the outcome to your decision? Outcome is dominated by other causes and noise Partly attributable with effort Clean and unambiguous attribution
Volume. How many times have you made this exact decision with a known outcome? Fewer than 20 Dozens Hundreds or more

Add the four scores for a total out of 8.

Total Verdict What to do with your gut
6 to 8 High-validity decision Trust trained intuition, especially under time pressure. Your fast judgment is likely carrying real information
3 to 5 Mixed Use intuition to generate options and a checklist or model to select between them. Never let the feel of certainty close the decision
0 to 2 Wicked decision Do not trust the gut, however experienced you are. Use base rates, explicit criteria set before you look, and outside opinions. This is where confident people are most often wrong

The uncomfortable output for most readers is that their highest-stakes decisions score lowest. Hiring a senior leader, entering a new market, choosing a strategic bet: low volume, slow feedback, contaminated attribution. These are precisely the decisions people most often make on instinct and most often defend with experience.

The scorecard applied to common modern decisions

The table below is our own scoring, offered as a worked example rather than as measurement. Score your own version; local conditions change the numbers.

Recurring decision Regularity Feedback speed Feedback clarity Volume Total Verdict
Judging whether a piece of written work is good 2 2 1 2 7 High validity
Estimating how long a familiar task will take you 2 2 2 2 8 High validity
Debugging a system you built and maintain 2 2 2 2 8 High validity
Reviewing an AI output in your own field of expertise 2 1 1 2 6 High validity
Deciding which candidate to hire for a senior role 1 0 0 0 1 Wicked
Predicting which product feature customers will adopt 1 1 1 1 4 Mixed
Choosing a career move 0 0 0 0 0 Wicked
Judging whether a new market is worth entering 0 0 0 0 0 Wicked

Notice the shape of the result. The decisions where your instinct is genuinely excellent tend to be the craft-level ones you perform constantly with visible results. The decisions the world rewards most highly are the ones where instinct is worth least. That gap is the entire argument for building judgment deliberately rather than accumulating years.

Building judgment where the environment will not teach you

If your important decisions score low, the answer is not resignation. It is to manufacture the conditions the environment fails to supply. Each of the following exists to convert a wicked decision into a slightly kinder one.

Write the prediction down before the outcome. The single highest-return habit available. Before a consequential decision, record what you expect to happen, by when, and with what confidence as a percentage. Wicked environments corrupt learning mostly through hindsight: without a record, you will remember having expected whatever happened. A prediction log is the minimum instrument that makes feedback honest. This is exactly the discipline the Good Judgment Project used to turn ordinary volunteers into forecasters who outperformed professional analysts.

Attach reasons, not just calls. Record why you decided. When the outcome arrives you need to know whether you were right for the right reason, because being right for the wrong reason teaches you a false cue and is worse than being wrong.

Shorten the loop artificially. If the real outcome takes three years, define an intermediate marker you can check in three months. It is a weaker signal than the real thing, and vastly better than nothing.

Start from base rates. In low-validity domains, the outside view beats the inside story. How often do acquisitions of this type work? How long do projects like this actually take? Begin with the reference class, then adjust for specifics, rather than building from the vivid particulars in front of you.

Fix your criteria before you look. Decide what would make a candidate or an opportunity good before you meet the candidate or see the opportunity. This is the practical defence against a coherent impression forming first and the reasons assembling themselves afterwards.

Be a fox. Tetlock’s finding was that forecasters who draw on many frameworks and update readily beat those committed to a single big idea. In practice: hold several explanations at once, and treat a change of mind as a working method rather than an embarrassment.

Seek disconfirmation from people who owe you nothing. Judgment fails quietly when everyone consulted has the same information and the same incentives.

The AI-era wrinkle

Two things change when a capable model sits inside your workflow, and they pull in opposite directions.

The first is genuinely good. AI shortens feedback loops. Work that once took a week to produce and a month to evaluate can now be drafted, tested and revised in an afternoon. More iterations with visible results is precisely what turns a wicked environment kinder, which means some domains are becoming more learnable than they were.

The second is a new failure mode: borrowed confidence. A fluent, well-organised, superbly reasoned answer produces the same feeling of certainty that genuine expertise produces, without the underlying validity. This is the Hogarth problem with a new delivery mechanism. Your scorecard is the defence. On a high-validity decision in your own craft, your reaction to a machine’s output is informative and you should trust your discomfort. On a wicked decision, your comfort with a confident answer tells you nothing at all, and the correct response is to check it against base rates and outside views rather than against your sense of how right it sounds.

This is also why taste and discernment have become trainable competitive assets rather than vague qualities, a case we develop in The Taste Stack. For the operating layer, the scorecard here slots directly into the routing logic in The Personal Decision Stack, and the frameworks worth having loaded before you decide are indexed in Mental Models That Actually Matter in 2026.

The short version

Judgment is not a personality trait and not a reward for tenure. It is what you get when a learnable environment meets deliberate feedback, and it stays local to the decisions where those conditions held. In every other decision, the honest move is to stop asking your gut and start building the scaffolding that makes a good decision possible without it. The people who get this right are not the ones with the most experience. They are the ones who know exactly which of their experience counts.

FAQ

So is intuition reliable or not?
It depends entirely on where it was trained. Kahneman and Klein concluded that trustworthy intuition requires a regular environment plus prolonged practice with good feedback. Meet both and expert intuition is excellent, often better than formal analysis under time pressure. Miss either and it is confidence without accuracy.

I have twenty years in my field. Does the scorecard really apply to me?
Twenty years in a high-validity domain is exactly what builds real intuition. Twenty years in a wicked one mostly builds certainty. That is why Shanteau found experienced clinicians and parole officers performing little better than novices while weather forecasters and chess masters performed superbly. Years are an input, not a result.

Does this mean I should stop trusting myself on big decisions?
It means the bigger the decision, the less you should rely on the feeling of knowing and the more you should rely on structure: recorded predictions, base rates, criteria set in advance, disconfirming input. Use your intuition to generate candidate answers. Do not let it be the judge.

What is the fastest way to start improving?
Keep a decision journal. Before each meaningful call, write your prediction, your confidence as a percentage, and your reasons. Review quarterly. It costs a few minutes per decision and is the only reliable way to find out whether your judgment is calibrated in a domain that will not tell you on its own.

Can AI improve my judgment or does it just outsource it?
Both are available and you choose which. Used to generate counterarguments, surface base rates, and stress-test a prediction you have already committed to in writing, it improves calibration. Used to hand you a confident conclusion you adopt because it reads well, it degrades your judgment while raising your confidence, which is the worst combination available.

Sources and further reading

  • Kahneman, D., and Klein, G. (2009). Conditions for Intuitive Expertise: A Failure to Disagree. American Psychologist, 64(6), 515-526.
  • Shanteau, J. (1992). Competence in Experts: The Role of Task Characteristics. Organizational Behavior and Human Decision Processes, 53(2).
  • Macnamara, B. N., Hambrick, D. Z., and Oswald, F. L. (2014). Deliberate Practice and Performance in Music, Games, Sports, Education, and Professions: A Meta-Analysis. Psychological Science.
  • Hogarth, R. M., Lejarraga, T., and Soyer, E. (2015). The Two Settings of Kind and Wicked Learning Environments. Current Directions in Psychological Science.
  • Hogarth, R. M. Educating Intuition, University of Chicago Press.
  • Tetlock, P. E. (2005). Expert Political Judgment: How Good Is It? How Can We Know? Princeton University Press.
  • Tetlock, P. E., and Gardner, D. Superforecasting: The Art and Science of Prediction, on the Good Judgment Project and the IARPA forecasting tournament.
  • Klein, G. Sources of Power: How People Make Decisions, MIT Press, for the naturalistic decision-making research programme.

This content was compiled with the support of AI following in-depth research, then written and prepared for publication by the CEOtudent editorial team.

This post is also available in: Türkçe Français Español Deutsch

Benzer içerikler