DijitalGelişim
0

From Tutorial Hell to Real Skill: The Project-Based Learning Ladder

TL;DR. Tutorial hell is what happens when a method that is right for beginners is kept past the point where it works. Worked examples help novices (a 2023 meta-analysis of 55 mathematics studies put the average effect at g = 0.48), but the advantage fades and then reverses as knowledge grows. The feeling of progress does not fade with it. In a Harvard physics experiment, students in active classes scored 0.46 standard deviations higher and felt they had learned 0.56 standard deviations less (Deslauriers and colleagues, 2019). Students who reread a text about 14 times recalled 40 percent of it a week later, against 61 percent for students who read it about 3 times and tested themselves (Roediger and Karpicke, 2006). The opposite mistake, jumping straight to an unguided project, is not supported either: the project-based programmes with positive trial results were heavily structured and led by teachers. AI adds a faster version of the same trap. Students given a plain chatbot raised their practice scores by 48 percent and then scored 17 percent worse on an unassisted exam (Bastani and colleagues, 2025), and developers learning a new library with AI help averaged 50 percent on a comprehension quiz against 67 percent without it (Shen and Tamkin, 2026). The Project-Based Learning Ladder, a CEOtudent framework, turns this evidence into six rungs with an exit test for each: study an example, complete one, modify one, rebuild it from memory, extend it to a new specification, then ship your own project and get feedback. The CEO move is to cap the tutorials you have in progress at one. The student move is to treat every exit test as the lesson.

This article is the practical companion to deliberate practice in the age of AI. That piece asks where your practice repetitions should go. This one gives the staircase that gets you from following along to building on your own.

What is tutorial hell, and why does it feel like progress?

“Tutorial hell” is a community term, not a research one. It describes a learner who finishes tutorial after tutorial, can follow each one successfully, and still cannot build anything without one. The term comes from programming, but the pattern is the same in design, data analysis, languages or music: fluent while guided, stuck when the guide is removed.

The reason it persists is that the experience feels like learning. Three experiments show how unreliable that feeling is.

Smooth delivery raises confidence, not recall. Carpenter, Wilford, Kornell and Mullaney (2013, Psychonomic Bulletin and Review) showed students the same short lecture delivered either fluently or haltingly by the same speaker. In the first experiment (42 students), those who watched the fluent version predicted they would remember far more (d = 1.03), but actual recall did not differ, so overconfidence was .22 after the fluent video against .03 after the halting one. A second experiment with 70 students again found no significant difference in test scores (.44 against .40).

Effort is misread as failure. Deslauriers and colleagues (2019, PNAS) taught 149 Harvard physics students the same material in two ways, with each student experiencing both. After active sessions students scored 0.46 standard deviations higher on a test of learning, yet rated their feeling of learning 0.56 standard deviations lower than after polished passive lectures. The authors’ explanation is that the confusion and effort of active work is taken as a sign of poor learning when it signals the opposite.

Rereading wins today and loses next week. Roediger and Karpicke (2006, Psychological Science) had students either study a passage repeatedly or study it once and then recall it. Five minutes later, repeated study was ahead (83 percent against 71 percent). One week later the order had reversed: 40 percent for repeated study against 61 percent for study followed by three recall tests. The repeated-study group had read the passage an average of 14.2 times, the testing group 3.4 times, and the repeated-study group was the more confident of the two.

A tutorial is a well-produced worked example with a fluent presenter. On all three counts it is built to feel effective.

Is watching tutorials a mistake?

No. For a beginner it is the right first step, and skipping it is its own error.

Kirschner, Sweller and Clark (2006, Educational Psychologist) reviewed the evidence on instruction with minimal guidance and concluded that novices learn less from it, and less efficiently, than from strongly guided instruction. Beginners do not yet have the knowledge in long-term memory that makes open-ended problem solving productive. The best-studied form of guidance is the worked example. A 2023 meta-analysis of 55 studies and 181 effect sizes in mathematics (Barbieri and colleagues, Educational Psychology Review) found a medium average effect of g = 0.48 for learning from worked examples. The same analysis found that examples of correct solutions worked better than incorrect or mixed ones, and that adding self-explanation prompts reduced the effect on average. Only its abstract was accessible to us.

The catch is in the second half of the same research programme. Kirschner and colleagues note that the advantage of guidance recedes once learners have enough prior knowledge to guide themselves, and that the worked-example effect first disappears and then reverses as expertise increases. Kalyuga, Ayres, Chandler and Sweller (2003) named this the expertise reversal effect: methods that work well for inexperienced learners can stop working, and can even backfire, once learners know more. Renkl and Atkinson (2003) drew the practical conclusion: examples early, problem solving later, with a fading procedure in between in which the learner takes over more of each solution step by step.

Read this way, tutorial hell is not a character flaw. It is an expertise-reversal problem (our reading of the research, not a term the authors use): the learner has outgrown the format and kept using it.

The format is also the default. In the 2025 Stack Overflow Developer Survey (49,009 responses; 33,454 answered this question), respondents who learned to code in the past year most often named technical documentation (67.8 percent), other online resources (58.7 percent), Stack Overflow itself (51.4 percent), videos (50 percent) and AI tools (44 percent). Among respondents aged 18 to 24, videos reached 55.2 percent. The sample is self-selected and mostly working developers (76.2 percent are professionals; 5.2 percent say they are learning to code), so it describes how developers keep learning rather than how beginners start.

What does the evidence say about doing versus watching?

Two large bodies of data point the same way, with different strengths.

In one Coursera psychology course with built-in interactive exercises, Koedinger and colleagues (2015) modelled what predicted quiz scores among the 939 students who took the final exam. An increase of one standard deviation in exercises done was associated with a 0.44 standard deviation gain; the same increase in videos watched or pages read was associated with about .065 each, which the authors summarise as more than six times the benefit for doing. The data are correlational and come from students who chose how to study, and only 4 percent of the 27,720 people who registered took the final at all.

The experimental evidence comes from classrooms. Freeman and colleagues (2014, PNAS) pooled 225 studies of undergraduate science, engineering and mathematics courses. Active learning raised exam performance by 0.47 standard deviations, and average failure rates were 21.8 percent under active learning against 33.8 percent under traditional lecturing, a gap of 12 percentage points.

And retrieval beats re-exposure even against a respectable study technique. Karpicke and Blunt (2011, Science) compared recalling a science text from memory with drawing concept maps of it. A week later the retrieval group scored 0.67 against 0.45. In the second experiment, 101 of 120 students (84 percent) did better after retrieval practice, although 90 of the 120 (75 percent) had predicted that concept mapping would be at least as good.

The fluency trap: what felt better and what tested better

The table is a CEOtudent synthesis. It lines up six experiments, four from before generative AI and two with it, by the same two questions: which condition looked or felt better at the time, and which did better on a later unassisted test.

Experiment Looked or felt better at the time Did better on the later test
Carpenter and colleagues, 2013 (42 students) Fluent lecture: much higher predicted recall (d = 1.03) Neither; actual recall did not differ
Deslauriers and colleagues, 2019 (149 students) Passive lecture: feeling of learning 0.56 SD higher Active class: 0.46 SD higher
Roediger and Karpicke, 2006 (180 students) Repeated study: 83 percent against 71 percent after 5 minutes, and higher confidence Repeated testing: 61 percent against 40 percent after 1 week
Karpicke and Blunt, 2011 (120 students) Concept mapping: 75 percent expected it to be as good or better Retrieval practice: 84 percent of students did better with it
Bastani and colleagues, 2025 (nearly 1,000 students) Plain chatbot: practice scores up 48 percent Control group: chatbot students scored 17 percent worse
Shen and Tamkin, 2026 (52 developers) AI assistance: task finished about two minutes sooner (not significant) Hand coding: 67 percent against 50 percent on the quiz

In every row the condition that felt or looked better in the moment was not the one that produced more learning. The repeated-study group in the 2006 experiment is the cleanest illustration: about four times the exposure (14.2 readings against 3.4) for roughly two thirds of the retention (40 percent against 61 percent). That is the arithmetic of tutorial hell.

Why does “just build a project” fail for beginners too?

Because the research that supports projects did not test unguided ones.

The strongest argument for struggling first is the literature on problem solving before instruction. Sinha and Kapur (2021, Review of Educational Research) pooled 53 studies with 166 comparisons and found a moderate advantage for attempting problems before being taught (g = 0.36, 95 percent confidence interval 0.20 to 0.51), rising to between 0.37 and 0.58 when the design was implemented faithfully. Two conditions matter. The struggle was always followed by instruction, and the pattern ran the other way for younger learners (second to fifth graders) and for domain-general skills, where instruction first did better.

The randomized trials of project-based learning tell the same story about structure. In 46 Michigan schools with 2,371 third graders, a project-based science curriculum raised standardized test scores by 0.277 standard deviations (Krajcik and colleagues, 2023); the intervention included four designed units, materials, professional learning for teachers and assessments after each unit. In a study of 3,645 students across 68 randomized schools, a project-based version of two Advanced Placement courses raised the share earning a qualifying exam score by 4 percentage points overall, and by 7.6 points among students who sat the exam, from 37.2 to 44.8 percent (Saavedra and colleagues, 2021). Teachers in that programme received a four-day summer institute, four further full days during the year and coaching on demand. The authors also caution that schools dropping out after randomization weakens the causal claim.

Projects work when they sit on a scaffold. A beginner who is told to “just build something” has been handed the last rung of a ladder with the lower rungs removed.

What changes when AI is in the loop?

AI can be the fastest tutorial ever made, or a tutor. The evidence separates the two sharply.

As an answer machine it deepens the trap. Bastani and colleagues (2025, PNAS) ran a field experiment with nearly 1,000 students in grades 9 to 11 at a high school in Turkey, across four 90-minute mathematics sessions. A plain GPT-4 chatbot raised scores on practice problems by 48 percent; a version instructed to give hints rather than answers raised them by 127 percent. On the later exam without AI, the plain-chatbot group scored 17 percent worse than students who had practised with textbooks only, while the hint-giving tutor removed the harm without producing a gain. Students in the plain-chatbot group did not perceive that they had done worse. The chatbot itself gave a correct answer only 51 percent of the time.

The same holds for working developers. Shen and Tamkin (2026, a preprint from Anthropic) asked 52 mostly junior software engineers to learn an unfamiliar Python library, randomly assigned to work with or without an AI assistant. The AI group averaged 50 percent on a comprehension quiz against 67 percent for the hand-coding group (d = 0.738, p = 0.010), a gap of 17 points, with the largest differences on debugging questions. The AI group finished about two minutes sooner, which was not statistically significant. How people used the assistant mattered: participants who delegated the code, leaned on the assistant more as the task went on, or used it to debug by iteration averaged below 40 percent, while those who asked conceptual questions, asked for explanations with the code, or generated code and then worked to understand it averaged 65 percent or more. These groups contain 2 to 7 people each, so the pattern is descriptive.

As a structured tutor it can help. In a randomized trial with 194 Harvard physics students (Kestin and colleagues, 2025, Scientific Reports), a carefully designed AI tutor produced a median post-test score of 4.5 against 3.5 for an active-learning class, in a median of 49 minutes. The same authors warn against using AI where students are likely to treat it as a crutch.

The rule that follows: on the ladder, AI may explain, hint, quiz and review. It may not produce the thing the rung exists to make you produce. For the wider evidence, see is AI making you worse at thinking.

The Project-Based Learning Ladder

The ladder is a CEOtudent editorial framework. It arranges the findings above into six rungs in order of decreasing guidance. It has not been tested as a package; each rung names the research it borrows from, and each has an exit test you must pass without help before moving up.

Rung What you do Evidence it borrows from Exit test Permitted AI use
1. Study Work through one complete example and label what each step is for Worked examples (g = 0.48 in mathematics); guidance for novices You can say in plain words what every step does and why Explaining steps
2. Complete Finish a partly built version: fill gaps or put shuffled steps in order Completion problems took a third less time than writing from scratch with similar learning; fading You complete it with the example closed Hints, not solutions
3. Modify Change a working version and predict the result before running it Predict-run-modify teaching in 13 schools Your predictions are usually right Checking a prediction after you make it
4. Rebuild Recreate the whole example from memory, days later Retrieval practice: 61 percent against 40 percent after a week It works with every reference closed None until you are done, then review
5. Extend Add a feature the example never covered, attempting it before looking anything up Problem solving before instruction (g = 0.36), followed by instruction The extension works and you can explain the part you had to look up Explanation after your own attempt
6. Ship Build something of your own for a real user and collect specific feedback Structured projects (0.277 SD; 4 points); high-information feedback (d = 0.99) A user or reviewer has responded and you have shipped a revision Review and critique of your work

Rung 1, study. One good example beats five skimmed ones. Margulieux, Morrison and Decker (2020) gave half of 265 students in an introductory Java course worked examples whose steps were labelled with their purpose. On the first quiz, 68 percent of that group reached the top two levels for explaining code in plain English, against 37 percent of the control group. The labelled group did better on quizzes but not on exams, and fewer of them dropped or failed the course. Label the steps yourself if the tutorial does not.

Rung 2, complete. Ericson, Margulieux and Rick (2017) had 135 students in a first Python course practise either by arranging mixed-up lines of code into a correct program, by fixing broken code, or by writing the code from scratch. The four practice problems took an average of 473 seconds as arrangement problems, 679 seconds as fixing and 714 seconds as writing, and there was no significant difference in learning immediately or one week later among the 82 who returned. The timing figures come from the first author’s dissertation chapter on the same study, and the authors note that the equivalence needs replication.

Rung 3, modify. The PRIMM approach (predict, run, investigate, modify, make) was evaluated in 13 schools with 493 pupils aged 11 to 14 over 8 to 12 weeks alongside a control group (Sentance, Waite and Kallia, 2019). The abstract reports better results for the PRIMM group on the post-test without giving an effect size, and the learners were school pupils, so treat this rung as the least firmly supported.

Rung 4, rebuild. This is the rung that tutorial hell skips, and the one with the strongest experimental support. Close everything and recreate the project. It will feel worse than rewatching. The retention data above say it is worth more.

Rung 5, extend. Try first, then look up exactly what blocked you. The order matters: the benefit in the research belongs to attempts that are followed by instruction, not to struggle left unresolved.

Rung 6, ship. Feedback is the point of shipping. Wisniewski, Zierer and Hattie (2020) pooled 435 studies and found an average effect of d = 0.48 for feedback on learning, but the type mattered more than the average: feedback rich in information reached 0.99, simple reinforcement or punishment 0.24, and 17 percent of all effects were negative. Ericsson, Krampe and Tesch-Romer (1993) made immediate, informative feedback a defining condition of deliberate practice. Ask reviewers what is wrong and how to fix it, not whether they like it. More on what practice does and does not buy is in the 10,000-hour myth.

How do you find your rung and run the ladder each week?

Find your rung with one test. Take the last tutorial you completed and try rung 4: rebuild it with everything closed. If you can, start at rung 5. If you cannot, drop to rung 2 or 3 with the same material instead of starting a new tutorial.

The CEO move: manage work in progress.

  1. One tutorial in progress. Do not start another until the current one has passed rung 4. This single constraint converts consumption into inventory you can actually use.
  2. Pick the project before the tutorial. Decide what rung 6 will be, then choose guidance that serves it. For choosing the skill itself, use the 4-factor decision matrix.
  3. Count exits, not hours. Track how many rungs you cleared this week, not how long you watched. Realistic timelines for specific skills are in time to competence for 20 skills.
  4. Buy feedback early. Line up the user, reviewer or community for rung 6 at the start, so shipping has somewhere to go.

The student move: make the test the lesson.

  • Predict before you run. Write down what you expect before executing, compiling, rendering or submitting anything.
  • Keep an error log. One line per mistake: what you expected, what happened, what you now know. It is your own high-information feedback.
  • Rebuild on a delay. Repeat rung 4 a week later for anything you want to keep. The schedule that prevents fading is in skill decay and maintenance, and the ranking of techniques behind it is in 12 study techniques ranked.
  • Use AI as the tutor, not the author. Ask it to explain, to quiz you and to critique what you built. A protocol for that is in the 20-hour AI tutor protocol.

When is the ladder the wrong tool?

  • When you are at the very beginning. With no vocabulary at all, stay on rungs 1 and 2 longer. Needing guidance at this stage is what the research predicts, not a weakness.
  • When mistakes are costly or irreversible. In clinical, safety-critical or financial work, the upper rungs belong under supervision, not in a solo project.
  • When the skill is domain-general or the learner is young. The struggle-first evidence reversed for early school grades and for general skills.
  • When you need the result, not the skill. If a task is a one-off, following a tutorial or delegating it is a reasonable decision. Make it deliberately, and do not count it as practice.

What does the verified data say? Key studies at a glance

Study Design and sample Reported result
Barbieri and colleagues, 2023 Meta-analysis, 55 studies, 181 effect sizes, mathematics Worked examples g = 0.48; self-explanation prompts lowered the effect
Deslauriers and colleagues, 2019 Randomized crossover, 149 physics students Active classes 0.46 SD higher on the test, 0.56 SD lower on feeling of learning
Roediger and Karpicke, 2006 Two experiments, 120 and 180 students After 1 week 56 against 42 percent (one test against restudy) and 61 against 40 percent (three tests against repeated study)
Karpicke and Blunt, 2011 Two experiments, 80 and 120 students Retrieval 0.67 against concept mapping 0.45; 84 percent did better with retrieval
Freeman and colleagues, 2014 Meta-analysis, 225 studies Exam performance up 0.47 SD; failure 21.8 against 33.8 percent
Ericson, Margulieux and Rick, 2017 Experiment, 135 students Practice time 473 seconds (arranging) against 714 (writing); no significant learning difference
Sinha and Kapur, 2021 Meta-analysis, 53 studies, 166 comparisons g = 0.36 for problem solving before instruction; reversed for grades 2 to 5
Krajcik and colleagues, 2023 Cluster randomized trial, 46 schools, 2,371 pupils 0.277 SD on a standardized science test
Saavedra and colleagues, 2021 Randomized schools, 3,645 students 4 points more earning a qualifying score; 7.6 points among exam-takers
Wisniewski, Zierer and Hattie, 2020 Meta-analysis, 435 studies Feedback d = 0.48; high-information 0.99; reinforcement or punishment 0.24
Bastani and colleagues, 2025 Field experiment, nearly 1,000 students Practice up 48 and 127 percent; unassisted exam 17 percent worse with the plain chatbot
Shen and Tamkin, 2026 Randomized experiment, 52 developers Quiz 50 against 67 percent (d = 0.738); no significant time saving

FAQ

What is tutorial hell?
It is the state of completing tutorials successfully while being unable to build without one. It is a community term. The research behind it concerns fluency illusions and the expertise reversal effect.

Are tutorials bad for learning?
No. Guided examples are the most efficient start for a novice, with an average effect of g = 0.48 in mathematics. They stop paying once you can follow them easily, which is the moment to move to completion, modification and rebuilding.

How many tutorials should I finish before starting a project?
The research offers no number. Use a test instead: if you can rebuild the last tutorial project from memory, you are ready to extend it and then to start your own.

Does project-based learning actually work?
Structured versions do. Randomized trials found gains of 0.277 standard deviations in elementary science and 4 percentage points in qualifying exam scores in two high-school courses. Both programmes came with designed materials and extensive teacher support.

Is it fine to use AI while learning to code?
It depends on the use. Students with a plain chatbot scored 17 percent worse on a later exam without it, and developers who delegated code to an assistant averaged below 40 percent on a comprehension quiz. Asking for explanations and hints was associated with scores of 65 percent or more.

Why does rebuilding from memory feel so much worse than rewatching?
Because effort is easy to mistake for failure. In the retention experiments, the method that felt less effective produced more learning a week later.

Sources

  1. Carpenter SK, Wilford MM, Kornell N, Mullaney KM. Appearances can be deceiving: instructor fluency increases perceptions of learning without increasing actual learning. Psychonomic Bulletin and Review. 2013;20:1350-1356.
  2. Deslauriers L, McCarty LS, Miller K, Callaghan K, Kestin G. Measuring actual learning versus feeling of learning in response to being actively engaged in the classroom. PNAS. 2019;116(39):19251-19257.
  3. Roediger HL, Karpicke JD. Test-enhanced learning: taking memory tests improves long-term retention. Psychological Science. 2006;17(3):249-255.
  4. Karpicke JD, Blunt JR. Retrieval practice produces more learning than elaborative studying with concept mapping. Science. 2011;331(6018):772-775.
  5. Kirschner PA, Sweller J, Clark RE. Why minimal guidance during instruction does not work. Educational Psychologist. 2006;41(2):75-86.
  6. Barbieri CA, Miller-Cotto D, Clerjuste SN, Chawla K. A meta-analysis of the worked examples effect on mathematics performance. Educational Psychology Review. 2023;35:11. Abstract.
  7. Kalyuga S, Ayres P, Chandler P, Sweller J. The expertise reversal effect. Educational Psychologist. 2003;38(1):23-31. Abstract.
  8. Renkl A, Atkinson RK. Structuring the transition from example study to problem solving in cognitive skill acquisition. Educational Psychologist. 2003;38(1):15-22. Abstract.
  9. Stack Overflow. 2025 Developer Survey, official results (Developers and Methodology pages).
  10. Koedinger KR, Kim J, Jia JZ, McLaughlin EA, Bier NL. Learning is not a spectator sport: doing is better than watching for learning from a MOOC. Proceedings of the Second ACM Conference on Learning at Scale. 2015:111-120.
  11. Freeman S, Eddy SL, McDonough M, Smith MK, Okoroafor N, Jordt H, Wenderoth MP. Active learning increases student performance in science, engineering, and mathematics. PNAS. 2014;111(23):8410-8415.
  12. Sinha T, Kapur M. When problem solving followed by instruction works: evidence for productive failure. Review of Educational Research. 2021;91(5):761-798. Abstract.
  13. Krajcik J, Schneider B, Miller EA, Chen I-C, Bradford L, Baker Q, Bartz K, Miller C, Li T, Codere S, Peek-Brown D. Assessing the effect of project-based learning on science learning in elementary schools. American Educational Research Journal. 2023;60(1):70-102. Abstract.
  14. Saavedra AR, Liu Y, Haderlein SK, Rapaport A, Garland M, Hoepfner D, Morgan KL, Hu A. Knowledge in Action efficacy study over two years. USC Dornsife Center for Economic and Social Research. 2021.
  15. Bastani H, Bastani O, Sungu A, Ge H, Kabakci O, Mariman R. Generative AI without guardrails can harm learning: evidence from high school mathematics. PNAS. 2025;122(26):e2422633122.
  16. Shen JH, Tamkin A. How AI impacts skill formation. Preprint, Anthropic, 2026.
  17. Kestin G, Miller K, Klales A, Milbourne T, Ponti G. AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports. 2025;15:17458.
  18. Margulieux LE, Morrison BB, Decker A. Reducing withdrawal and failure rates in introductory programming with subgoal labeled worked examples. International Journal of STEM Education. 2020;7:19.
  19. Ericson BJ, Margulieux LE, Rick J. Solving Parsons problems versus fixing and writing code. Proceedings of the 17th Koli Calling International Conference on Computing Education Research. 2017:20-29. Abstract, with timing data from Ericson BJ, doctoral dissertation, Georgia Institute of Technology, 2018.
  20. Sentance S, Waite J, Kallia M. Teaching computer programming with PRIMM: a sociocultural perspective. Computer Science Education. 2019;29(2-3):136-176. Abstract.
  21. Wisniewski B, Zierer K, Hattie J. The power of feedback revisited: a meta-analysis of educational feedback research. Frontiers in Psychology. 2020;10:3087.
  22. Ericsson KA, Krampe RT, Tesch-Romer C. The role of deliberate practice in the acquisition of expert performance. Psychological Review. 1993;100(3):363-406.

The fluency-trap table and the Project-Based Learning Ladder are CEOtudent analyses built on the sources above. All other figures are reported as printed in those sources.


This content was compiled with the support of AI following in-depth research, then written and prepared for publication by the CEOtudent editorial team.

This post is also available in: Türkçe Français Español Deutsch

Benzer içerikler