GelişimStrateji
0

How to Learn From AI Mistakes: Turning Model Errors Into Your Own Skill Gains

A woman at a sunlit desk reviewing and correcting handwritten notes beside an open laptop

TL;DR. AI errors are frequent enough in real work to be a steady supply of practice material, and most people throw that material away. In a randomized field experiment with nearly a thousand high school maths students (Bastani and colleagues, PNAS, 2025), a plain GPT-4 interface raised practice scores by 48 percent, yet those students scored 17 percent lower than the control group once the tool was removed. The same tool gave a correct answer only 51 percent of the time. A Stanford and Yale evaluation found that leading legal research tools hallucinate between 17 and 33 percent of the time. On the learning side, a meta-analysis of 24 studies of error management training found a positive mean effect of d = 0.44, rising to d = 0.80 for transfer to structurally new tasks. The CEO move: own an AI Error Log and a written verification policy, so errors become a managed risk instead of a surprise. The student move: treat every caught error as a practice rep, classify it, rework it without the tool and review the log every week.

This piece belongs to the evaluation cluster. Start with the hub, the evaluation skill: judging AI output, for the overall method. For the mechanics behind the errors, see why AI hallucinates and how to catch model errors early. For deciding how much checking a task deserves, see trust calibration: when to trust AI recommendations.

Why treat an AI mistake as training data instead of a nuisance?

The usual reaction to a model error is to fix it and move on. That handles the output and wastes the lesson. Two bodies of research suggest a better use.

The first is the psychology of learning from errors. Janet Metcalfe’s review in the Annual Review of Psychology (2017) summarises the laboratory evidence in one sentence: “errorful learning followed by corrective feedback is beneficial to learning.” She adds a condition that matters for anyone working with AI: “Corrective feedback, including analysis of the reasoning leading up to the mistake, is crucial.” Fixing a wrong number is not enough. The gain comes from working out why it was wrong.

The second is the evidence on what happens when people use AI without that analysis. The authors of the PNAS field experiment state the risk plainly. Because generative AI is fallible, “users must vigilantly check its outputs and fix any issues present; if they fail to learn the underlying skills, then they may lack the expertise required to do so.” The skill of checking depends on the skill of doing. If AI use erodes the second, the first goes with it.

Every wrong answer from a model is either a small tax paid and forgotten, or a problem with feedback attached. That is the CEO and student lens applied to one object. A chief executive does not treat a defect as bad luck; the defect is logged, classified and traced to a process. A good student does not treat a wrong answer as an embarrassment; the wrong answer is the most informative item on the page. The AI Error Log described below is both at once.

How often is AI actually wrong?

The honest answer is that there is no single rate. Published error rates differ by more than an order of magnitude depending on the task. The figures below are as printed in their sources.

Task and system tested Published figure Source
Summarising a supplied news document, best models on the HHEM leaderboard Lowest hallucination rate 1.3% (two models tied), then 1.4% and 1.5% Stanford HAI, AI Index Report 2025
Short fact-seeking questions (SimpleQA, over 4,000 questions) Best-performing model answered 42.7% of questions successfully Stanford HAI, AI Index Report 2025
High school maths practice problems, GPT-4 behind a plain chat interface Correct answer 51% of the time; logical errors 42%; arithmetic errors 8% Bastani et al., PNAS, 2025
Legal research questions, commercial tools with retrieval (202 queries) Hallucination between 17% and 33% of the time Magesh et al., 2024 preprint
Direct, verifiable questions about federal court cases, general-purpose models Hallucination between 58% (ChatGPT 4) and 88% (Llama 2) of the time Dahl et al., 2024

Sources: Stanford Institute for Human-Centered AI, AI Index Report 2025, Chapter 3; Bastani et al., PNAS 122, 2025; Magesh et al., arXiv preprint 2405.20362, 2024; Dahl et al., arXiv 2401.01301, forthcoming in the Journal of Legal Analysis. All figures describe the specific models and dates tested and are not current rates for any product.

Three points follow from the table.

First, the rate depends on the task far more than on the brand. The Harvard Business School and Boston Consulting Group experiment gave this unevenness a name. Its abstract describes a “jagged technology frontier”, where AI assistance “improves performance for some tasks but worsens it for others, even within the same knowledge workflow and with a seemingly similar level of difficulty.”

Second, specialised tools reduce errors without removing them. The legal study tested products whose vendors had advertised an end to hallucinations. The authors report that Lexis+ AI answered 65 percent of queries accurately, that Westlaw’s AI-Assisted Research was accurate 42 percent of the time, and that “Over 1 in 6” queries led Lexis+ AI and Ask Practical Law AI to respond with misleading or false information. Their definition is useful for any error log: a response counts as hallucinated “if it is either incorrect or misgrounded”.

Third, the models are poor judges of their own errors. Dahl and colleagues conclude that the models “struggle to predict their own hallucinations, and often uncritically accept users’ incorrect legal assumptions.” A confident tone is not evidence.

What happens when people do not learn from AI errors?

The PNAS experiment is the clearest evidence, because it measured both assisted performance and unassisted performance afterwards. The study ran in a high school in Turkey in the fall semester of the 2023-2024 academic year, across four 90-minute sessions and about fifty classes. Students practised with one of two GPT-4 tools or with textbooks only, then sat a closed-book exam.

Measure Control (textbooks only) GPT Base (plain chat interface) GPT Tutor (teacher-designed prompts)
Assisted practice score, change against control (out of 1) Mean 0.28 +0.137 +0.361
Practice improvement as printed – 48% 127%
Unassisted exam, change against control (out of 1) – -0.054 -0.004
Exam effect as printed – 17% reduction Statistically indistinguishable from control

Source: Bastani H, Bastani O, Sungu A, Ge H, Kabakcı Ö, Mariman R, PNAS 122, e2422633122, 2025.

Two details make this more than a warning about schoolchildren. The authors audited the plain tool by asking it each of the 57 practice problems ten times. It “gives a correct answer only 51% of the time on average”, with logical errors 42 percent of the time and arithmetic errors 8 percent of the time. Students were practising with a tutor that was wrong about half the time, and the unassisted exam shows they did not turn those errors into learning.

They also did not notice. Students in the GPT Base arm “did not perceive that they performed worse or learned less.” The feeling of progress and the fact of progress came apart.

The consulting experiment shows the same pattern in professional work. Among 758 knowledge workers, those using GPT-4 on 18 tasks inside the frontier completed 12.2 percent more tasks and finished them 25.1 percent more quickly. On a task selected to be outside the frontier, they were “19% less likely to produce correct solutions compared with those without AI.” The tool did not announce which side of the frontier a task was on. Only the people who checked could know.

Older research on decision support systems explains the mechanism. A systematic review of automation bias in the Journal of the American Medical Informatics Association screened 13,821 papers and included 74. In its pooled analysis of four studies, the risk ratio was 1.26 (95% CI 1.11 to 1.44): when the system’s advice was wrong, it “increased the risk of an incorrect decision being made by 26%.” The review describes two kinds of human error: commission, following incorrect advice, and omission, failing to act because the system did not prompt it. Both belong in an error log, because both are yours, not the model’s.

What does learning science say about errors that are corrected?

If uncorrected errors harm, the research on corrected errors points the other way. The figures are as printed in each source.

Finding Evidence base Effect size as printed
Error management training, overall (Keith and Frese, 2008) 24 studies, N = 2,183 Cohen’s d = 0.44
Same, post-training transfer Moderator analysis d = 0.56
Same, structurally distinct tasks (adaptive transfer) Moderator analysis d = 0.80
Problem solving before instruction (Sinha and Kapur, 2021) 53 studies, 166 comparisons Hedge’s g 0.36, 95% CI 0.20 to 0.51
Same, high fidelity to productive failure principles Subset Hedge’s g between 0.37 and 0.58

Sources: Keith N, Frese M, Journal of Applied Psychology 93(1), 2008; Sinha T, Kapur M, Review of Educational Research 91(5), 2021. Both read at abstract level only.

Four findings translate directly into practice with AI.

Errors help most on new problems. Keith and Frese found the largest effect for adaptive transfer, 0.80 against 0.44 overall, about 1.8 times the average. Error-based training is “better suited than error-avoidant training methods for promotion of transfer to novel tasks.” Novel tasks are what an AI-era job increasingly consists of.

Trying first matters, even when the attempt fails. Kornell, Hays and Bjork ran six experiments in which participants had to guess answers they could not know before seeing them. Their conclusion: “Unsuccessful retrieval attempts enhanced learning with both types of materials.” The implication for AI use is a sequencing rule. Form your own answer, estimate or outline before reading the model’s.

Confident errors are the most correctable. Butterfield and Metcalfe expected high-confidence errors to be the hardest to fix. They found the opposite: “highly confident errors were the most likely to be corrected in a subsequent retest.” This is why the log below records your confidence before the check. The entries where you were sure and wrong are the most valuable ones.

Feedback has to include the reasoning. Metcalfe’s condition, quoted above, rules out the lazy version of an error log. A list of wrong outputs teaches little. A list of wrong outputs with the cause and the check that would have caught it is a curriculum.

One caution: Sinha and Kapur report that for “the learning of domain-general skills”, effect sizes favoured instruction first. Checking AI output is partly a general skill, so learn the checks explicitly as well. The taxonomy below is that instruction.

Which human skill does each type of AI error train?

CEOtudent editorial framework: the AI error taxonomy. The first six rows are model-side errors; the last two are human-side errors, taken from the automation bias literature. The mapping to skills is an editorial judgment, not a measured result.

Code Error type What it looks like Human skill it trains The rep
F Fabricated source or fact A citation, case, quote or statistic that does not exist Source verification Open the original and find the exact line before using it
M Misgrounded claim A real source cited for something it does not say Close reading Compare the claim with the cited passage word by word
L Logical error Plausible steps that do not follow, or the wrong method Reasoning and problem structuring Redo the first two steps unaided, then find where the paths split
A Arithmetic or unit error Right method, wrong number Estimation Write an order-of-magnitude estimate before reading the answer
P Accepted false premise The model builds on a wrong assumption in the prompt Question framing Ask the same question with the premise stated neutrally
S Scope drift or incomplete answer Answers a nearby question, or omits a required part Brief writing Rewrite the brief with a definition of done and rerun
C Commission (human) You accepted wrong output without checking Verification discipline Name the check that was skipped and add it to the policy
O Omission (human) You missed something because the model did not raise it Independent scanning List what a good answer must contain before prompting

The codes follow distinctions made in the sources: incorrect against misgrounded responses in the legal study, logical against arithmetic errors in the PNAS audit (42 against 8 percent), accepted false premises in Dahl and colleagues, and commission against omission in the automation bias review.

A pattern in the codes is a diagnosis. Many S entries mean the briefs are weak. Many C entries mean the verification policy is not being followed. Many F and M entries in one domain mean the tool is past its frontier for that task. For the brief-writing side of this loop, see the manager-of-AI playbook.

What should an AI Error Log contain?

CEOtudent editorial framework: the AI Error Log template. One row per caught error. A spreadsheet or a notes file is enough.

Column What to write Why it is there
Date and task What was being produced, and for whom Shows which workflows generate errors
Stakes Low, medium or high Links the entry to the verification policy
Tool and mode Model, with or without search or documents Error rates differ by task and setup
The error, quoted The exact wrong sentence or number Prevents vague memory
Code F, M, L, A, P, S, C or O Makes the weekly count possible
How it was caught Which check, or who found it downstream Shows which checks earn their time
Confidence before the check 0 to 100 that the output was right Surfaces confident misses, the most correctable kind
Cause on the human side Weak brief, missing context, skipped check, unfamiliar domain Moves the lesson from the model to the user
Correct answer and source The fix, with where it was verified The corrective feedback itself
Rule change The line added to the verification policy, or “none” Turns the error into a process change
Retest date When to rework this item unaided Schedules the practice rep

Two columns do most of the work. Entries caught downstream by a colleague, client or reviewer are escapes; their share of all entries shows whether the checking is working. The confidence column guards against the perception gap in the PNAS study: a number written before the check cannot be revised afterwards.

For the same before-and-after format applied to choices, see the decision journal template and protocol.

How much practice material is there?

CEOtudent analysis of the published rates above. The table converts each printed rate into what a person who checks every output would meet. It assumes each output is an independent draw at the printed rate, which is a simplification: real errors cluster by task type.

Context and printed rate Expected errors per 20 outputs Chance of at least one error in 10 outputs One error on average every Errors met in 50 weeks at 20 checked outputs a week
Document summaries, best models, 1.3% 0.3 12.3% 76.9 outputs 13
Legal research tools, low end, 17% 3.4 84.5% 5.9 outputs 170
Legal research tools, high end, 33% 6.6 98.2% 3.0 outputs 330
Plain GPT-4 maths answers, 49% (100 minus 51) 9.8 99.9% 2.0 outputs 490
General chatbot on case-law questions, low end, 58% 11.6 above 99.9% 1.7 outputs 580

Source: CEOtudent calculations (calc.py) from rates printed in the AI Index Report 2025, Magesh et al. 2024, Bastani et al. 2025 and Dahl et al. 2024. Illustrative, not a forecast for any current tool.

At a 17 percent rate, someone who accepts ten outputs unchecked has an 84.5 percent chance that at least one is wrong. The same rate, with checking, yields about 170 worked examples with feedback in a year. The general chatbot’s low end is 3.4 times that of the specialised tools, so tool choice changes the load without removing it.

At 1.3 percent an error arrives about once in 77 outputs, the condition under which attention drifts. Low-rate tasks need a fixed sampling rule in the policy, not vigilance.

How do you run the weekly review?

CEOtudent editorial protocol: 20 minutes, once a week. Over 50 weeks this is about 16.7 hours.

  1. Count and classify (5 minutes). Tally the week’s entries by code. Note the escapes, the entries someone else caught.
  2. Rework one error unaided (5 minutes). Take the entry with the highest confidence before the check. Without the tool, produce the correct answer or the correct method. This is the retrieval attempt; it works even when it fails, provided the correct answer is reviewed afterwards.
  3. Trace the cause (5 minutes). For the two most frequent codes, write one sentence on the human-side cause. Use Metcalfe’s condition as the test: the note must describe the reasoning that led to the miss, not only the miss.
  4. Change one rule (3 minutes). Add, tighten or delete one line of the verification policy. One change a week is enough; more will not be followed.
  5. Schedule the retest (2 minutes). Set a date to rework this week’s item again, and rework the item that came due from an earlier week.

Once a month, check that the share of escapes is falling and that the mix of codes is shifting away from C and S.

What does the CEO side look like: the verification policy?

A log without a policy is a diary. The policy is a short written document that says how much checking each class of work gets before it leaves your hands. Automation bias research supports writing it down. The JAMIA review lists “training and emphasizing user accountability” among the mitigators, and reports one study in which people who perceived themselves to be accountable made fewer automation bias errors. Parasuraman and Manzey are less optimistic, concluding that automation bias “cannot be prevented by training or instructions.” The two reviews agree on the practical point: good intentions are not a control. A named owner and a fixed procedure are.

CEOtudent editorial framework: a three-tier verification policy.

Tier Typical work Minimum check
1. Low stakes, reversible Drafts, brainstorming, internal notes Read once; sample-check one fact per document
2. Medium stakes Client-facing text, analysis that informs a decision, code that will run Verify every number, name, date and citation against a source; rerun calculations independently
3. High stakes or outside your expertise Legal, medical, financial or safety content; anything signed or published Tier 2 checks plus review by a qualified person; no source, no claim

Two rules keep the policy honest. Every escape moves that task type up one tier until a month passes without another. And every rule names its check, not an intention (“open the cited page”, not “be careful”).

The survey by Lee and colleagues at Microsoft Research and Carnegie Mellon shows why the policy cannot rest on feeling. Among 319 knowledge workers who shared 936 examples of using generative AI at work, “higher confidence in GenAI is associated with less critical thinking, while higher self-confidence is associated with more critical thinking.” Trust in the tool lowered checking exactly where checking was the only safeguard. The gap between tool confidence and self-confidence is the subject of the discernment gap.

Are you learning from AI mistakes? An 8-question scorecard

CEOtudent editorial scorecard. Score each question 0 (no), 1 (sometimes) or 2 (consistently). Maximum 16.

  1. Do you form your own estimate, outline or answer before reading the model’s?
  2. Do you record caught errors somewhere, with the exact wrong text?
  3. Do you classify each error by type?
  4. Do you write down your confidence before checking?
  5. Do you note the human-side cause, not only the model’s fault?
  6. Do you rework at least one error unaided each week?
  7. Do you have a written verification policy with tiers by stakes?
  8. Do you track errors that someone else caught after you?
Score Reading Next step
0-6 Errors are being fixed and forgotten Start the log with five columns: task, error, code, cause, fix
7-11 Errors are captured, not yet converted Add the confidence column and the weekly unaided rework
12-16 Errors are a working curriculum Track the monthly escape share and prune rules that never fire

What this does not mean

AI errors are not good, and more errors are not better. The learning research is about corrected errors in low-stakes settings. Metcalfe’s review recommends allowing errors “while they are in low-stakes learning situations”, not in work that reaches a client or a court.

The link from AI errors to human skill is an inference. The learning studies cited here examine people’s own errors, mostly in classrooms, laboratories and training courses. No study opened for this article tested whether logging a model’s errors improves the user’s judgment. The framework applies established principles to a new setting; it is not a measured effect.

The error rates are snapshots. They describe specific models tested in 2023 and 2024. The legal study’s authors caution that their estimate “is not meant to be an unbiased estimate” of the population rate of hallucinations in legal AI queries. Current tools may do better or worse on any given task.

The practice-material table is arithmetic, not prediction. It assumes independent errors at a constant rate and that every error is caught.

Frequently asked questions

How many errors are needed before a log is useful?
A handful. The weekly protocol works with one entry, because step 2 needs only one item to rework. Patterns in the codes usually need a few weeks of entries.

Should errors be logged when the stakes were trivial?
Yes, briefly. Low-stakes errors are the safest practice material, and they reveal the same weaknesses in briefs and checks that cause high-stakes errors.

Is it worth logging errors that were the user’s fault?
Those are the most useful entries. Codes C and O exist because the automation bias literature treats following wrong advice and missing unprompted issues as human errors.

Does a better model make the log unnecessary?
No. Lower error rates make each error rarer and therefore easier to miss. The log then shifts from a practice source to a sampling record that shows the checks are still being done.

What is the fastest first step?
Before the next AI answer on something that matters, write a one-line estimate of what the answer should be. Then compare. That single habit applies the retrieval finding and makes the first log entry easy to write.

Sources

  1. Bastani H, Bastani O, Sungu A, Ge H, Kabakcı Ö, Mariman R. Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences 122, e2422633122, 2025. A 2025 correction concerns an author affiliation only.
  2. Magesh V, Surani F, Dahl M, Suzgun M, Manning CD, Ho DE. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv preprint 2405.20362, 2024.
  3. Dahl M, Magesh V, Suzgun M, Ho DE. Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models. arXiv 2401.01301, 2024, forthcoming in the Journal of Legal Analysis.
  4. Maslej N, et al. The AI Index 2025 Annual Report. AI Index Steering Committee, Institute for Human-Centered AI, Stanford University, April 2025. Chapter 3, Responsible AI.
  5. Dell’Acqua F, McFowland E, Mollick E, Lifshitz-Assaf H, Kellogg KC, et al. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality. Organization Science 37(2): 403-423, 2026. Abstract only.
  6. Goddard K, Roudsari A, Wyatt JC. Automation bias: a systematic review of frequency, effect mediators, and mitigators. Journal of the American Medical Informatics Association 19(1): 121-127, 2012.
  7. Parasuraman R, Manzey DH. Complacency and Bias in Human Use of Automation: An Attentional Integration. Human Factors 52(3): 381-410, 2010. Abstract only.
  8. Lee HP, Sarkar A, Tankelevitch L, et al. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers. CHI Conference on Human Factors in Computing Systems, 2025.
  9. Metcalfe J. Learning from Errors. Annual Review of Psychology 68: 465-489, 2017. Abstract only.
  10. Keith N, Frese M. Effectiveness of Error Management Training: A Meta-Analysis. Journal of Applied Psychology 93(1): 59-69, 2008. Abstract only.
  11. Sinha T, Kapur M. When Problem Solving Followed by Instruction Works: Evidence for Productive Failure. Review of Educational Research 91(5): 761-798, 2021. Abstract only.
  12. Kornell N, Hays MJ, Bjork RA. Unsuccessful Retrieval Attempts Enhance Subsequent Learning. Journal of Experimental Psychology: Learning, Memory, and Cognition 35(4): 989-998, 2009.
  13. Butterfield B, Metcalfe J. Errors Committed With High Confidence Are Hypercorrected. Journal of Experimental Psychology: Learning, Memory, and Cognition 27(6): 1491-1494, 2001. Abstract only.

Data tables report figures as printed in their sources. The error taxonomy, Error Log template, practice-material table, weekly review protocol, verification policy and scorecard are CEOtudent analyses or editorial frameworks; the 1.8 and 3.4 multiples are CEOtudent calculations. Not used: the working paper version of the consulting experiment, which could not be opened; the legal study’s chart-only values for GPT-4; the 0.87 bias-adjusted estimate in Sinha and Kapur; and popular claims that chatbots are wrong a fixed share of the time, which no source opened here supports.


This content was compiled with the support of AI following in-depth research, then written and prepared for publication by the CEOtudent editorial team.

This post is also available in: Türkçe Français Español Deutsch

Benzer içerikler