{"id":325437,"date":"2026-08-26T11:30:00","date_gmt":"2026-08-26T08:30:00","guid":{"rendered":"https:\/\/ceotudent.com\/socratic-prompting-method-make-ai-teach-you"},"modified":"2026-08-26T11:30:00","modified_gmt":"2026-08-26T08:30:00","slug":"socratic-prompting-method-make-ai-teach-you","status":"publish","type":"post","link":"https:\/\/ceotudent.com\/en\/socratic-prompting-method-make-ai-teach-you","title":{"rendered":"The Socratic Prompting Method: How to Make AI Teach You Instead of Doing the Work for You"},"content":{"rendered":"<p><strong>TL;DR:<\/strong> There are now two large randomised trials that point in apparently opposite directions, and the gap between them is the entire practical lesson. Bastani and colleagues, publishing in PNAS in 2025, gave nearly a thousand Turkish high school maths students access to GPT-4 during practice. Performance during practice rose 48 percent for a plain ChatGPT-style interface and 127 percent for a version with learning safeguards. When access was removed for the exam, the plain-interface students scored 17 percent worse than students who never had access, while for the safeguarded group the authors report that the negative effect was largely mitigated, with no exam advantage over control. Kestin and colleagues, in Scientific Reports in 2025, ran a randomised trial in a Harvard physics course (N = 194) and found an AI tutor produced median learning gains more than double those of an expertly run active-learning classroom, in a median of 49 minutes against a 60-minute lesson. Our derived reading: the effective tutor delivered roughly 2.3 times the learning gain in 18.3 percent less time, which is about 2.9 times the gain per minute. What separates the two results is not the model. It is whether the interaction forces the learner to generate, or lets them receive. Below: the method, seven paste-ready prompt templates, a substitution ledger for the requests you already make, and a table of what the evidence supports and what it does not.<\/p>\n<p>You have almost certainly done this. You hit something you do not understand, you paste it into an assistant, you get back a clear explanation, and you feel the click of understanding.<\/p>\n<p>That feeling is not evidence of learning. It is one of the best-documented illusions in cognitive science, and it has a name.<\/p>\n<p>Rozenblit and Keil demonstrated it across twelve studies published in Cognitive Science in 2002. People &ldquo;feel they understand complex phenomena with far greater precision, coherence, and depth than they really do&rdquo;, and the illusion is specifically strongest for <em>explanatory<\/em> knowledge, which is exactly the kind of knowledge an AI explanation delivers. The effect is most robust, they found, where the environment supports real-time explanations with visible mechanisms. A chat window with an instant, fluent, well-structured answer is close to a purpose-built machine for producing that illusion.<\/p>\n<p>So the question is not whether AI can teach you. It clearly can. The question is what the interaction has to look like before the teaching actually lands.<\/p>\n<div id=\"ez-toc-container\" class=\"ez-toc-v2_0_84 counter-hierarchy ez-toc-counter ez-toc-grey ez-toc-container-direction\">\n<div class=\"ez-toc-title-container\">\n<p class=\"ez-toc-title\" style=\"cursor:inherit\">Table of Contents<\/p>\n<span class=\"ez-toc-title-toggle\"><a href=\"#\" class=\"ez-toc-pull-right ez-toc-btn ez-toc-btn-xs ez-toc-btn-default ez-toc-toggle\" aria-label=\"Toggle Table of Content\"><span class=\"ez-toc-js-icon-con\"><span class=\"\"><span class=\"eztoc-hide\" style=\"display:none;\">Toggle<\/span><span class=\"ez-toc-icon-toggle-span\"><svg style=\"fill: #999;color:#999\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" class=\"list-377408\" width=\"20px\" height=\"20px\" viewBox=\"0 0 24 24\" fill=\"none\"><path d=\"M6 6H4v2h2V6zm14 0H8v2h12V6zM4 11h2v2H4v-2zm16 0H8v2h12v-2zM4 16h2v2H4v-2zm16 0H8v2h12v-2z\" fill=\"currentColor\"><\/path><\/svg><svg style=\"fill: #999;color:#999\" class=\"arrow-unsorted-368013\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"10px\" height=\"10px\" viewBox=\"0 0 24 24\" version=\"1.2\" baseProfile=\"tiny\"><path d=\"M18.2 9.3l-6.2-6.3-6.2 6.3c-.2.2-.3.4-.3.7s.1.5.3.7c.2.2.4.3.7.3h11c.3 0 .5-.1.7-.3.2-.2.3-.5.3-.7s-.1-.5-.3-.7zM5.8 14.7l6.2 6.3 6.2-6.3c.2-.2.3-.5.3-.7s-.1-.5-.3-.7c-.2-.2-.4-.3-.7-.3h-11c-.3 0-.5.1-.7.3-.2.2-.3.5-.3.7s.1.5.3.7z\"\/><\/svg><\/span><\/span><\/span><\/a><\/span><\/div>\n<nav><ul class='ez-toc-list ez-toc-list-level-1 ' ><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-1\" href=\"https:\/\/ceotudent.com\/en\/socratic-prompting-method-make-ai-teach-you\/#The-two-trials-and-what-actually-separates-them\" >The two trials, and what actually separates them<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-2\" href=\"https:\/\/ceotudent.com\/en\/socratic-prompting-method-make-ai-teach-you\/#Why-generating-beats-receiving\" >Why generating beats receiving<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-3\" href=\"https:\/\/ceotudent.com\/en\/socratic-prompting-method-make-ai-teach-you\/#The-method\" >The method<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-4\" href=\"https:\/\/ceotudent.com\/en\/socratic-prompting-method-make-ai-teach-you\/#Seven-paste-ready-templates\" >Seven paste-ready templates<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-5\" href=\"https:\/\/ceotudent.com\/en\/socratic-prompting-method-make-ai-teach-you\/#The-substitution-ledger\" >The substitution ledger<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-6\" href=\"https:\/\/ceotudent.com\/en\/socratic-prompting-method-make-ai-teach-you\/#What-the-evidence-supports-and-what-it-does-not\" >What the evidence supports, and what it does not<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-7\" href=\"https:\/\/ceotudent.com\/en\/socratic-prompting-method-make-ai-teach-you\/#Fitting-it-into-how-you-actually-study\" >Fitting it into how you actually study<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-8\" href=\"https:\/\/ceotudent.com\/en\/socratic-prompting-method-make-ai-teach-you\/#The-CEO-and-the-student\" >The CEO and the student<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-9\" href=\"https:\/\/ceotudent.com\/en\/socratic-prompting-method-make-ai-teach-you\/#Frequently-asked-questions\" >Frequently asked questions<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-10\" href=\"https:\/\/ceotudent.com\/en\/socratic-prompting-method-make-ai-teach-you\/#Sources\" >Sources<\/a><\/li><\/ul><\/nav><\/div>\n<h2 id=\"the-two-trials-and-what-actually-separates-them\"><span class=\"ez-toc-section\" id=\"The-two-trials-and-what-actually-separates-them\"><\/span>The two trials, and what actually separates them<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Start with the harm case, because it is the more useful one.<\/p>\n<p>Bastani, Bastani, Sungu, Ge, Kabakc\u0131 and Mariman published &ldquo;Generative AI without guardrails can harm learning: Evidence from high school mathematics&rdquo; in the Proceedings of the National Academy of Sciences in July 2025. They ran a field experiment with nearly a thousand high school maths students across three conditions: GPT Base, a standard ChatGPT-style interface; GPT Tutor, the same underlying model with prompts designed to safeguard learning; and a control group with textbook and notes only.<\/p>\n<p>During the assisted practice sessions, GPT Base students scored 48 percent higher than control and GPT Tutor students scored 127 percent higher. Then the assistance was removed and everyone sat the exam. GPT Base students scored 17 percent <em>lower<\/em> than the control group. In the authors&rsquo; framing, &ldquo;unfettered access to GPT-4 can harm educational outcomes&rdquo;, because &ldquo;students attempt to use GPT-4 as a crutch during practice problem sessions, and subsequently perform worse on their own.&rdquo; The safeguards in GPT Tutor largely eliminated that negative effect.<\/p>\n<p>Now do the arithmetic that most coverage skips.<\/p>\n<table>\n<thead>\n<tr>\n<th>Condition<\/th>\n<th>Gain during practice<\/th>\n<th>Effect on the unassisted exam<\/th>\n<th>Share of the practice gain that survived<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>GPT Base<\/td>\n<td>+48%<\/td>\n<td>-17% vs control<\/td>\n<td>-0.35<\/td>\n<\/tr>\n<tr>\n<td>GPT Tutor<\/td>\n<td>+127%<\/td>\n<td>No measurable advantage over control; negative effect largely mitigated<\/td>\n<td>~0<\/td>\n<\/tr>\n<tr>\n<td>Control<\/td>\n<td>baseline<\/td>\n<td>baseline<\/td>\n<td>not applicable<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p><em>Ratio column is a CEOtudent derivation: exam effect divided by practice gain. It expresses how much of the visible performance boost persisted once the tool was removed.<\/em><\/p>\n<p>Neither AI condition produced a measurable exam advantage. The prompt-level guardrails did something valuable but modest: they converted a 17 percent loss into a wash. If your mental model was &ldquo;a good tutor prompt makes AI a learning accelerator&rdquo;, this trial does not support it. It supports &ldquo;a good tutor prompt stops AI from being a learning <em>decelerator<\/em>.&rdquo;<\/p>\n<p>Now the gain case.<\/p>\n<p>Kestin, Miller, Klales, Milbourne and Ponti published &ldquo;AI tutoring outperforms in-class active learning&rdquo; in Scientific Reports in June 2025. In a Harvard undergraduate physics course (N = 194), students alternated between an expertly run active-learning class and a purpose-built AI tutor. Median post-test score was 4.5 in the AI condition (N = 142) against 3.5 in the classroom condition (N = 174), from a combined pre-test median of 2.75 (N = 316). A Mann-Whitney rank-sum test on the post-score distributions gave z = -5.6, p &lt; 10 to the minus 8. Median time on task in the AI condition was 49 minutes, against a 60-minute in-class lesson.<\/p>\n<p>Working those medians through, which is our own calculation and not one the paper reports:<\/p>\n<table>\n<thead>\n<tr>\n<th>Measure<\/th>\n<th>AI tutor<\/th>\n<th>Active-learning class<\/th>\n<th>Ratio<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Median post-test<\/td>\n<td>4.5<\/td>\n<td>3.5<\/td>\n<td>&#8211;<\/td>\n<\/tr>\n<tr>\n<td>Median gain over the 2.75 pre-test baseline<\/td>\n<td>1.75<\/td>\n<td>0.75<\/td>\n<td>2.33x<\/td>\n<\/tr>\n<tr>\n<td>Median time on task<\/td>\n<td>49 min<\/td>\n<td>60 min<\/td>\n<td>18.3% less<\/td>\n<\/tr>\n<tr>\n<td>Gain per minute<\/td>\n<td>0.0357<\/td>\n<td>0.0125<\/td>\n<td>2.86x<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p><em>CEOtudent derivation from the published medians. Medians do not decompose cleanly, so treat the per-minute figure as an order-of-magnitude comparison rather than a precise effect size. The authors themselves note a ceiling effect in the post-test scores.<\/em><\/p>\n<p>So: same underlying model family, same year, one trial finds harm and one finds a roughly 2.3-fold gain. The difference is in the design, and the Kestin paper documents it explicitly.<\/p>\n<p>Their tutor was built against seven pedagogical practices: fostering active engagement, managing cognitive load, promoting a growth mindset, scaffolding content, ensuring accuracy of information and feedback, delivering feedback in a targeted and timely fashion, and allowing self-pacing. Crucially, they report that a system prompt alone &ldquo;could not reliably provide enough structure to scaffold problems with multiple parts&rdquo;, so the <em>platform<\/em> enforced sequential progression through each part of each problem. They also supplied the model with detailed step-by-step solutions rather than letting it generate them, precisely because next-token prediction is unreliable on complex maths. Eighty-three percent of students rated the tutor&rsquo;s explanations as good as or better than those from human instructors.<\/p>\n<p>That is the honest headline. <strong>Prompt-level guardrails prevent harm. Structural design produces gain.<\/strong> You cannot get the Kestin result out of a single clever prompt, and anyone selling you one is overselling. What you can do is get a good deal closer than the default, and that is what the method below is for.<\/p>\n<h2 id=\"why-generating-beats-receiving\"><span class=\"ez-toc-section\" id=\"Why-generating-beats-receiving\"><\/span>Why generating beats receiving<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>The mechanism is not mysterious, and it predates AI by decades.<\/p>\n<p>Roediger and Karpicke showed in Psychological Science in 2006 that taking a memory test does not merely assess knowledge, it strengthens it. Students who studied prose passages and then took recall tests without feedback outperformed students who simply restudied the same material the same number of times, on delayed tests at two days and one week. Notably, on an immediate five-minute test the restudy group did better, and the restudy group was also <em>more confident<\/em> they would remember. Short-term fluency and long-term retention pull in opposite directions.<\/p>\n<p>That is the same trap as the explanation illusion, wearing different clothes. Reading a good explanation feels like the five-minute test. Producing the explanation yourself feels worse and works better.<\/p>\n<p>VanLehn&rsquo;s 2011 review in Educational Psychologist put numbers on how much tutoring can buy. Against no tutoring, human tutoring came in at an effect size of d = 0.79 and intelligent tutoring systems at d = 0.76, close enough that the human advantage largely disappears when the system operates at step-level rather than answer-level granularity. That last clause is the operative one for prompting. <strong>Answer-level interaction underperforms. Step-level interaction is where the effect lives.<\/strong><\/p>\n<p>There is also a preliminary neural signal worth noting with appropriate caution. Kosmyna and colleagues at MIT Media Lab circulated a preprint in 2025, not peer-reviewed, in which 54 participants wrote essays across three sessions in one of three conditions: LLM, search engine, or no tools. EEG showed the brain-only group with the strongest and most distributed connectivity, search-engine users intermediate, and LLM users the weakest, with cognitive activity scaling down in proportion to external tool use. Self-reported ownership of the essays was lowest in the LLM group, and the authors report that LLM users &ldquo;struggled to accurately quote their own work.&rdquo; Treat it as suggestive rather than settled: it is a preprint with a small sample. But it points the same direction as the behavioural evidence.<\/p>\n<h2 id=\"the-method\"><span class=\"ez-toc-section\" id=\"The-method\"><\/span>The method<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>The Socratic prompting method has one rule and five moves. The rule is:<\/p>\n<blockquote>\n<p><strong>Never let the model complete a step you could attempt.<\/strong><\/p>\n<\/blockquote>\n<p>Everything else is machinery for enforcing that rule against your own impatience.<\/p>\n<p><strong>Move 1: Declare the contract before you ask anything.<\/strong> The default behaviour of every assistant is to be maximally helpful, which means answering. You have to override that once, at the top, and then hold it.<\/p>\n<p><strong>Move 2: Front-load your own attempt.<\/strong> Say what you think first, however wrong. This is the generation step, and it is the one people skip. A wrong attempt that gets corrected produces more durable learning than a right answer you read.<\/p>\n<p><strong>Move 3: Work at step granularity.<\/strong> One step, one exchange. This is VanLehn&rsquo;s finding operationalised, and it is the thing the Kestin team had to enforce at the platform level because a prompt alone would not hold it.<\/p>\n<p><strong>Move 4: Close every concept with a recall test you did not write.<\/strong> Ask the model to test you later, from memory, with the material hidden. This is Roediger and Karpicke made practical.<\/p>\n<p><strong>Move 5: Force an explanation with the source closed.<\/strong> The explanation illusion collapses the moment you have to produce the explanation without looking. That collapse is the diagnostic.<\/p>\n<h2 id=\"seven-paste-ready-templates\"><span class=\"ez-toc-section\" id=\"Seven-paste-ready-templates\"><\/span>Seven paste-ready templates<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>These are the working tools. Each one implements one part of the method. Adapt the bracketed parts.<\/p>\n<p><strong>1. The opening contract.<\/strong> Paste this once at the start of a learning session and refer back to it whenever the model drifts.<\/p>\n<blockquote>\n<p>For this entire conversation you are my tutor, not my assistant. Rules: never give me a complete answer or a finished solution. Work one step at a time. Before each step, ask me what I think comes next and wait for my reply. If I am wrong, do not correct me directly; ask a question that makes the error visible to me. If I ask you to just tell me, remind me of this rule once, then ask whether I want to override it. Keep your turns under 120 words. Begin by asking me what I already know about [TOPIC].<\/p>\n<\/blockquote>\n<p><strong>2. The attempt-first frame.<\/strong> For any problem you would normally paste in cold.<\/p>\n<blockquote>\n<p>Here is a problem: [PROBLEM]. Here is my attempt, which may be wrong: [YOUR ATTEMPT, EVEN IF PARTIAL]. Do not solve it. Tell me only whether my first step is sound, and if not, ask me one question that would help me find the error myself.<\/p>\n<\/blockquote>\n<p><strong>3. The step gate.<\/strong> When the model runs ahead.<\/p>\n<blockquote>\n<p>Stop. You just gave me [N] steps. Discard everything after step one. Ask me what step two should be and wait.<\/p>\n<\/blockquote>\n<p><strong>4. The Feynman check.<\/strong> The illusion detector. Use it after you feel you understand.<\/p>\n<blockquote>\n<p>I am going to explain [CONCEPT] to you from memory, without looking anything up. Do not help me while I do it. When I finish, list only the places where my explanation was vague, circular, or missing a mechanism. Do not fill the gaps in; just name them.<\/p>\n<\/blockquote>\n<p><strong>5. The delayed retrieval test.<\/strong> The single highest-value template on this list, and the least used.<\/p>\n<blockquote>\n<p>Do not explain anything now. Write me five questions on [TOPIC] that I should be able to answer from memory tomorrow. Include one question that requires applying the idea to a situation we never discussed. Hold the answers until I have attempted all five.<\/p>\n<\/blockquote>\n<p><strong>6. The steelman inversion.<\/strong> For judgment-type material rather than procedural material.<\/p>\n<blockquote>\n<p>I believe [YOUR POSITION] about [TOPIC]. Do not agree or disagree. Ask me the three questions that someone who holds the opposite view would ask, one at a time. Wait for each answer before asking the next.<\/p>\n<\/blockquote>\n<p><strong>7. The transfer probe.<\/strong> The exam-condition simulator, which is precisely what the Bastani students never faced during practice.<\/p>\n<blockquote>\n<p>Give me one problem on [TOPIC] that is structurally similar to what we just worked through but superficially different, so surface pattern-matching will not solve it. Do not give hints. I will attempt it and paste my full reasoning; only then respond.<\/p>\n<\/blockquote>\n<h2 id=\"the-substitution-ledger\"><span class=\"ez-toc-section\" id=\"The-substitution-ledger\"><\/span>The substitution ledger<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>This table is a CEOtudent editorial framework. It maps the requests people actually type against a tutor-mode rewrite, and names the learning mechanism each rewrite recruits. The mechanisms are drawn from the published literature cited in this piece; the mapping is ours.<\/p>\n<table>\n<thead>\n<tr>\n<th>What you usually type<\/th>\n<th>Answer-mode outcome<\/th>\n<th>Tutor-mode rewrite<\/th>\n<th>Mechanism recruited<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>&ldquo;Explain [concept] to me&rdquo;<\/td>\n<td>Fluent explanation, high confidence, low retention<\/td>\n<td>&ldquo;Ask me what I already believe about [concept], then correct only what is wrong&rdquo;<\/td>\n<td>Generation before feedback<\/td>\n<\/tr>\n<tr>\n<td>&ldquo;Solve this problem&rdquo;<\/td>\n<td>Correct solution, zero transfer<\/td>\n<td>&ldquo;Here is my first step. Is it sound? One question only.&rdquo;<\/td>\n<td>Step-level tutoring<\/td>\n<\/tr>\n<tr>\n<td>&ldquo;Summarise this article&rdquo;<\/td>\n<td>Usable summary, no encoding<\/td>\n<td>&ldquo;I will summarise it. Then tell me what I missed.&rdquo;<\/td>\n<td>Retrieval practice<\/td>\n<\/tr>\n<tr>\n<td>&ldquo;Write the code for X&rdquo;<\/td>\n<td>Working code you cannot debug<\/td>\n<td>&ldquo;Give me the function signature and the first line. I will write the rest, then you review.&rdquo;<\/td>\n<td>Scaffolded production<\/td>\n<\/tr>\n<tr>\n<td>&ldquo;Is this argument any good?&rdquo;<\/td>\n<td>An evaluation you inherit<\/td>\n<td>&ldquo;Ask me the three questions a critic would ask.&rdquo;<\/td>\n<td>Self-explanation<\/td>\n<\/tr>\n<tr>\n<td>&ldquo;Give me a study plan&rdquo;<\/td>\n<td>A plan you will not follow<\/td>\n<td>&ldquo;Write me five recall questions for tomorrow, answers withheld.&rdquo;<\/td>\n<td>Spaced retrieval<\/td>\n<\/tr>\n<tr>\n<td>&ldquo;What are the main points?&rdquo;<\/td>\n<td>A list you will forget<\/td>\n<td>&ldquo;I will list them from memory. Name only the gaps.&rdquo;<\/td>\n<td>Illusion detection<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>The pattern across the right-hand column is consistent: every rewrite moves the cognitive work from the model back to you, and reserves the model for the thing it is genuinely better at than a textbook, which is targeted, immediate, self-paced feedback on your specific error. That is exactly the capability VanLehn identified as the source of the tutoring effect, and the two capabilities the Kestin team singled out as impossible to deliver in a classroom.<\/p>\n<h2 id=\"what-the-evidence-supports-and-what-it-does-not\"><span class=\"ez-toc-section\" id=\"What-the-evidence-supports-and-what-it-does-not\"><\/span>What the evidence supports, and what it does not<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>This table is a CEOtudent editorial framework. It grades common claims about AI and learning against the strongest published test of each, using a consistent rule: <strong>Supported<\/strong> means a randomised or large field study directly tested it; <strong>Partial<\/strong> means the evidence points that way but the design does not isolate the claim; <strong>Unsupported<\/strong> means the strongest test contradicts it.<\/p>\n<table>\n<thead>\n<tr>\n<th>Claim<\/th>\n<th>Strongest test<\/th>\n<th>Result<\/th>\n<th>Verdict<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Using AI during practice improves your practice performance<\/td>\n<td>Bastani et al., PNAS 2025, ~1,000 students<\/td>\n<td>+48% base, +127% tutor<\/td>\n<td>Supported<\/td>\n<\/tr>\n<tr>\n<td>Improved practice performance means you learned more<\/td>\n<td>Same trial, exam with tool removed<\/td>\n<td>-17% for base condition<\/td>\n<td>Unsupported<\/td>\n<\/tr>\n<tr>\n<td>A good tutor prompt turns AI into a learning accelerator<\/td>\n<td>Same trial, GPT Tutor condition<\/td>\n<td>No exam advantage over control reported<\/td>\n<td>Unsupported<\/td>\n<\/tr>\n<tr>\n<td>A good tutor prompt prevents AI from harming learning<\/td>\n<td>Same trial, GPT Tutor vs GPT Base<\/td>\n<td>Negative effect largely mitigated<\/td>\n<td>Supported<\/td>\n<\/tr>\n<tr>\n<td>A well-designed AI tutor can beat expert classroom teaching<\/td>\n<td>Kestin et al., Scientific Reports 2025, N = 194<\/td>\n<td>Median gains more than double, less time<\/td>\n<td>Supported<\/td>\n<\/tr>\n<tr>\n<td>Structured AI tutoring always beats classroom teaching<\/td>\n<td>Same paper, authors&rsquo; own caveat<\/td>\n<td>Not presumed for complex synthesis and higher-order tasks<\/td>\n<td>Unsupported<\/td>\n<\/tr>\n<tr>\n<td>Step-level interaction matters more than answer-level<\/td>\n<td>VanLehn, Educational Psychologist 2011<\/td>\n<td>d = 0.76 for step-based systems vs d = 0.79 human tutoring<\/td>\n<td>Supported<\/td>\n<\/tr>\n<tr>\n<td>Being tested beats rereading for long-term retention<\/td>\n<td>Roediger and Karpicke, Psychological Science 2006<\/td>\n<td>Testing wins at two days and one week; restudy wins at five minutes<\/td>\n<td>Supported<\/td>\n<\/tr>\n<tr>\n<td>Feeling you understand indicates you understand<\/td>\n<td>Rozenblit and Keil, Cognitive Science 2002<\/td>\n<td>Illusion of explanatory depth across twelve studies<\/td>\n<td>Unsupported<\/td>\n<\/tr>\n<tr>\n<td>Heavy LLM use reduces cognitive engagement while writing<\/td>\n<td>Kosmyna et al., preprint 2025, N = 54<\/td>\n<td>Weakest EEG connectivity, lowest ownership<\/td>\n<td>Partial<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>Read the middle rows together. The most expensive mistake in AI-assisted learning is not using the tool. It is trusting the practice-session signal. Your performance while the tool is open tells you almost nothing about your performance when it is closed, and in the one trial that measured both, the correlation ran the wrong way.<\/p>\n<h2 id=\"fitting-it-into-how-you-actually-study\"><span class=\"ez-toc-section\" id=\"Fitting-it-into-how-you-actually-study\"><\/span>Fitting it into how you actually study<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>The method costs time. That is the point, and it is also the reason people abandon it by the third session. Three ways to make it stick.<\/p>\n<p><strong>Split the session explicitly.<\/strong> Assistant mode and tutor mode should not share a window. Use the assistant to gather material, find sources, and clear away logistics. Then open a separate conversation, paste the opening contract, and do the learning there. Mixing them means the answering reflex leaks into the learning half.<\/p>\n<p><strong>Budget by intent, not by topic.<\/strong> Not everything is worth learning. If you will never need to produce it unaided, let the model produce it. Reserve tutor mode for the small set of things where your unassisted capability is the actual asset. The framework for making that call is in <a href=\"https:\/\/ceotudent.com\/en\/the-evaluation-skill-judging-ai-output\">the evaluation skill<\/a>, and the wider boundary question in <a href=\"https:\/\/ceotudent.com\/en\/should-you-let-ai-make-the-decision-delegation-boundaries\">what to delegate and what never to automate<\/a>.<\/p>\n<p><strong>Test on the schedule, not on the feeling.<\/strong> The delayed retrieval template is the whole method in one prompt. Fire it at the end of every session for the next day, and answer it before you open any source. If you cannot, you did not learn it, whatever the session felt like.<\/p>\n<p>For the fuller apparatus around this: <a href=\"https:\/\/ceotudent.com\/en\/what-the-evidence-says-about-learning-12-study-techniques-ranked\">what the evidence says about learning<\/a> ranks the study techniques themselves, <a href=\"https:\/\/ceotudent.com\/en\/ai-tutors-2026-which-tools-teach-which-give-answers\">AI tutors 2026<\/a> compares which products are built to teach rather than to answer, <a href=\"https:\/\/ceotudent.com\/en\/how-to-learn-anything-20-hours-ai-tutor-protocol\">the 20-hour AI tutor protocol<\/a> is the compressed sprint version, <a href=\"https:\/\/ceotudent.com\/en\/deliberate-practice-in-the-age-of-ai\">deliberate practice in the age of AI<\/a> covers the repetition layer, and <a href=\"https:\/\/ceotudent.com\/en\/how-to-build-a-personal-curriculum-with-ai-90-days\">building a personal curriculum with AI<\/a> sets the ninety-day structure this method runs inside.<\/p>\n<h2 id=\"the-ceo-and-the-student\"><span class=\"ez-toc-section\" id=\"The-CEO-and-the-student\"><\/span>The CEO and the student<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>This is the cleanest expression of the two-sided lens on the whole site.<\/p>\n<p>A CEO delegates aggressively and is right to. Work that does not need your hands should not have your hands on it, and refusing to delegate is not diligence, it is a bottleneck you built yourself. Most AI use should be exactly this: hand it over, check the output, move on.<\/p>\n<p>A student refuses to delegate the one thing that is theirs. Not the output, the capability. Because the capability is what determines whether you can check the output at all, which is the entire reason the Bastani authors flag long-term productivity as the thing at risk: generative AI &ldquo;is fallible and users must check its outputs&rdquo;, and you cannot check what you never learned.<\/p>\n<p>The Socratic method is how you hold both positions at once. Delegate the work. Refuse to delegate the learning. The prompt templates above are just the mechanics of drawing that line every day, in a tool whose default setting is to erase it.<\/p>\n<h2 id=\"frequently-asked-questions\"><span class=\"ez-toc-section\" id=\"Frequently-asked-questions\"><\/span>Frequently asked questions<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p><strong>Does using AI actually make you learn less?<\/strong><br \/>\nIt depends entirely on the interaction design. In the PNAS trial, students with an unrestricted ChatGPT-style interface during practice scored 17 percent worse on an unassisted exam than students who never had access. For students using a version with learning safeguards, the authors report the negative effect was largely mitigated, with no exam advantage over control. So unrestricted use produced measurable harm; safeguarded use did not.<\/p>\n<p><strong>What is Socratic prompting?<\/strong><br \/>\nA way of interacting with an AI model in which you never let it complete a step you could attempt yourself. You state a contract up front, submit your own attempt before asking anything, work one step per exchange, and end every session with a delayed recall test you did not write.<\/p>\n<p><strong>Can a prompt really turn ChatGPT into a good tutor?<\/strong><br \/>\nPartially. The PNAS trial found that prompt-level safeguards largely mitigated a 17 percent learning loss, but no exam advantage over the control group was reported. The Harvard trial that did produce large gains used platform-level sequencing and pre-supplied step-by-step solutions, not a system prompt alone. Expect a prompt to prevent harm reliably and to produce gains inconsistently.<\/p>\n<p><strong>Why does the explanation feel like understanding when it is not?<\/strong><br \/>\nBecause of the illusion of explanatory depth, documented by Rozenblit and Keil across twelve studies. People consistently overestimate how precisely and coherently they understand mechanisms, and the illusion is strongest exactly for explanatory knowledge and in environments that supply fluent real-time explanations. The test is producing the explanation yourself with the source closed.<\/p>\n<p><strong>How long should a Socratic session be?<\/strong><br \/>\nThe Harvard trial&rsquo;s median time on task was 49 minutes against a 60-minute classroom lesson, with 70 percent of students finishing under an hour. That is a reasonable ceiling for one concept. The binding constraint is not duration but whether you generated before you received.<\/p>\n<p><strong>Is it worth doing this for everything I learn?<\/strong><br \/>\nNo. Reserve it for material where your unassisted capability is the asset, meaning things you will later need to produce, judge, or defend without the tool open. For everything else, let the model do the work and check the output.<\/p>\n<h2 id=\"sources\"><span class=\"ez-toc-section\" id=\"Sources\"><\/span>Sources<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<ul>\n<li>Bastani, Bastani, Sungu, Ge, Kabakc\u0131 and Mariman, Generative AI without guardrails can harm learning: Evidence from high school mathematics, Proceedings of the National Academy of Sciences, volume 122, issue 26, 2025<\/li>\n<li>Kestin, Miller, Klales, Milbourne and Ponti, AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting, Scientific Reports, volume 15, article 17458, 2025<\/li>\n<li>VanLehn, The Relative Effectiveness of Human Tutoring, Intelligent Tutoring Systems, and Other Tutoring Systems, Educational Psychologist, volume 46, issue 4, 2011<\/li>\n<li>Roediger and Karpicke, Test-enhanced learning: taking memory tests improves long-term retention, Psychological Science, volume 17, issue 3, 2006<\/li>\n<li>Rozenblit and Keil, The misunderstood limits of folk science: an illusion of explanatory depth, Cognitive Science, volume 26, issue 5, 2002<\/li>\n<li>Kosmyna and colleagues, Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task, arXiv preprint, 2025, not peer-reviewed<\/li>\n<\/ul>\n<hr>\n<p><em>This content was compiled with the support of AI following in-depth research, then written and prepared for publication by the CEOtudent editorial team.<\/em><\/p>\n","protected":false},"excerpt":{"rendered":"<p>The largest field experiment on AI and learning gave nearly a thousand high school students GPT-4 during maths practice. Grades during practice rose 48 percent. Then the tool was taken away and the same students scored 17 percent worse on the exam than students who never had it. A guardrailed tutor version lifted practice performance 127 percent and left exam performance with no measurable advantage over the control group. Read that carefully: the safeguarded design did not create a learning advantage, it prevented a learning loss. Meanwhile a Harvard physics trial found an AI tutor more than doubled median learning gains against an active-learning classroom, in 49 median minutes against 60. The difference between those two results is not the model. It is the structure of the interaction. This piece turns that structure into a reusable prompting method with ready-to-paste templates, a substitution ledger showing how to rewrite the requests you already make, and an honest table of what the evidence does and does not support.<\/p>\n","protected":false},"author":1,"featured_media":325442,"comment_status":"open","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[4599,5],"tags":[],"class_list":["post-325437","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-gelisim","category-is"],"_links":{"self":[{"href":"https:\/\/ceotudent.com\/en\/wp-json\/wp\/v2\/posts\/325437","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/ceotudent.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/ceotudent.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/ceotudent.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/ceotudent.com\/en\/wp-json\/wp\/v2\/comments?post=325437"}],"version-history":[{"count":0,"href":"https:\/\/ceotudent.com\/en\/wp-json\/wp\/v2\/posts\/325437\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/ceotudent.com\/en\/wp-json\/wp\/v2\/media\/325442"}],"wp:attachment":[{"href":"https:\/\/ceotudent.com\/en\/wp-json\/wp\/v2\/media?parent=325437"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/ceotudent.com\/en\/wp-json\/wp\/v2\/categories?post=325437"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/ceotudent.com\/en\/wp-json\/wp\/v2\/tags?post=325437"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}