Gelişimİş
0

Fast Upskilling Protocols: What the Research Says vs. What Actually Works in Practice

Professional practising recall from memory beside a sunlit window with a closed book and review cards

TL;DR. The learning-science literature is unusually clear about what works. A 2021 meta-analysis covering 242 studies, 1,619 effects and 169,179 unique participants puts distributed practice first at d = 0.85 and practice testing second at d = 0.74, against an overall mean of 0.56. The European Continuing Vocational Training Survey shows what employers actually deploy: in 2020, 54.9% of EU enterprises ran training courses and 29.4% sent people to conferences, while job rotation and secondments reached 12.7% and learning circles 13.4%. The formats that structurally produce spacing and retrieval are the ones almost nobody runs. Three things then break the simple advice. The interleaving effect reverses sign on verbal material, favouring blocked practice at g = -0.39. The best-measured techniques were mostly tested on surface and factual outcomes, not on the deep relational skills adult upskilling targets. And the effects were much larger for lower-ability than higher-ability learners, which is the opposite of the population buying upskilling. The protocol below takes all three constraints seriously.

Two datasets that have never been put side by side

Anyone can find the list of evidence-based study techniques. Our own ranking of 12 study techniques by research strength sets out that hierarchy, and it is not in dispute.

What is missing from every version of that advice is the second half of the question. Knowing that distributed practice outperforms re-reading tells you what to do alone at a desk. It tells you nothing about why the training you actually receive at work looks the way it does, or which of the formats on offer is worth your time.

There is a public dataset for that, and it is rarely used in learning writing. Eurostat runs the Continuing Vocational Training Survey across EU enterprises and publishes, by format, the share of enterprises providing each kind of training. Put it against the effect-size table and the gap becomes visible.

Verified data: measured effect by technique, and deployment share by format

Learning technique (Donoghue and Hattie, 2021) Utility class Cases Unique participants Effect size d Nearest enterprise format (Eurostat CVTS) Share of EU enterprises, 2020
Distributed practice High 150 152,952 0.85 Job rotation, exchanges or secondments 12.7%
Practice testing High 374 6,033 0.74 Guided on-the-job training 43.1%
Elaborative interrogation Moderate 254 2,138 0.56 Learning or quality circles 13.4%
Imagery Moderate 135 1,052 0.56 not separately surveyed not reported
Self explanation Moderate 93 804 0.54 Self-directed learning 29.1%
Mnemonics Low 107 580 0.50 not separately surveyed not reported
Re-reading Low 113 1,529 0.47 not separately surveyed not reported
Interleaved practice Low 104 972 0.47 not separately surveyed not reported
Underlining Low 56 1,129 0.44 not separately surveyed not reported
Summarization Low 234 1,990 0.44 CVT courses and conference attendance 54.9% and 29.4%

Effect sizes, utility classes, case counts and participant counts are quoted directly from Table 1 of Donoghue and Hattie (2021), overall mean d = 0.56. Deployment shares are quoted from the Eurostat Continuing Vocational Training Survey, percentage of all EU enterprises providing each type of training, 2020 reference year. The mapping between a technique and an enterprise format is a CEOtudent editorial judgement about which format structurally produces that technique, not a statistical link; the two datasets share no common unit and are not joined numerically. Rows marked “not separately surveyed” have no corresponding category in the survey and are left empty rather than assigned a proxy.

Verified data: what EU enterprises actually provide, 2005 to 2020

Training format 2005 2010 2015 2020
Any continuing vocational training 55.6% 63.6% 70.5% 67.4%
CVT courses 46.5% 55.0% 60.2% 54.9%
CVT courses, external 41.8% 48.3% 54.1% 46.8%
CVT courses, internal 23.9% 29.5% 34.3% 35.0%
Guided on-the-job training 27.5% 30.7% 41.2% 43.1%
Conferences, workshops, trade fairs and lectures 29.7% 32.7% 37.2% 29.4%
Self-directed learning 9.9% 12.5% 19.5% 29.1%
Learning or quality circles 9.1% 8.8% 11.8% 13.4%
Job rotation, exchanges or secondments 8.5% 8.8% 11.6% 12.7%
No continuing vocational training 44.4% 36.4% 29.5% 32.6%

Eurostat, Continuing Vocational Training Survey, enterprises providing training by type of training, percentage of all enterprises, EU 27 countries from 2020, all size classes. Categories overlap; an enterprise providing several formats appears in several rows.

CEOtudent editorial framework: reading the two together

Three observations, none of which either dataset states on its own.

The highest-scoring technique is attached to the least-deployed format. Distributed practice measured d = 0.85, the top of the table, and the only enterprise format that structurally forces repeated encounters with the same material over months is job rotation and secondment, at 12.7% of enterprises. The course, which by construction compresses exposure into a block, reaches 54.9%.

The fastest-growing format moved the responsibility onto the individual. Self-directed learning went from 9.9% of enterprises in 2005 to 29.1% in 2020, roughly tripling and growing faster than any other category in the series. That is the single most consequential line in the table for a reader of this publication. The protocol is no longer designed by anybody but you.

The formats that did not grow are the ones that build spacing in. Learning circles went from 9.1% to 13.4% across fifteen years and job rotation from 8.5% to 12.7%. Both remain marginal. The structural conditions the evidence favours are not being supplied by employers at scale, and the trend does not suggest they will be.

This is the CEO-and-student split with a measurement attached. The student half asks which technique has the better effect size. The CEO half asks who is responsible for building the system that delivers it, looks at a survey showing that answer shifting onto the individual, and schedules accordingly.

Where the headline advice breaks

Three constraints that the popular version of this literature routinely drops. All three come from the same papers that produce the encouraging numbers.

Interleaving reverses sign on verbal material. Brunmair and Richter’s meta-analysis of 59 studies, 238 effect sizes and 158 samples found a moderate overall interleaving effect at Hedges’ g = 0.42, strongest for paintings at g = 0.67 and weaker for mathematical tasks at g = 0.34. But for studies based on words it found an advantage for blocking, at g = -0.39, and results for expository texts were ambiguous with non-significant overall effects. Interleaving is not a universal instruction. On vocabulary, terminology and definitions, which is a large share of what adult upskilling actually involves, the evidence points the other way.

The outcomes measured were mostly surface ones. Donoghue and Hattie state directly that the majority of studies in their meta-analysis were based on surface or factual outcomes, and that caution is needed when applying the findings to deeper and more relational outcomes. Nearly all professional upskilling is deep and relational: judgement, integration, transfer to messy cases. The d = 0.85 is real, and it was not measured on the thing you are trying to learn.

The effects were larger for lower-ability learners. The same paper reports that effects were much greater for lower than higher ability students, and that the presence of feedback and near versus far transfer were both important moderators. An experienced professional learning an adjacent skill sits at the end of the distribution where the measured lift is smaller.

None of this makes the techniques wrong. It makes the numbers ceilings rather than forecasts, and it explains why people who follow the advice faithfully often get less than the headline suggests.

What the spacing research actually specifies

The most useful and least quoted finding in this literature is about timing, and it is quantitative.

Cepeda and colleagues taught more than 1,350 people a set of facts, gave them a review after a gap of up to 3.5 months, and tested them again after a further delay of up to a year. At any given test delay, increasing the gap between study sessions first improved and then gradually reduced final performance. The optimal gap grew as the test delay grew. But measured as a proportion of the test delay, the optimal gap fell from about 20 to 40% of a one-week delay to about 5 to 10% of a one-year delay. The authors conclude that the interaction implies many educational practices are highly inefficient.

Their earlier review, covering 839 assessments across 317 experiments in 184 articles, reached the same structural point: the interstudy interval and the retention interval operate jointly, so no single spacing rule can be correct for every horizon.

Most popular summaries get this backwards. They report the shrinking percentage and conclude that longer horizons need relatively tighter review. In absolute time, the opposite is true, and absolute time is what goes in a calendar.

CEOtudent derived schedule: review gaps implied by the published ratios

How long you need to retain it Optimal gap as a share of that horizon Gap in days, derived
1 week (7 days) 20% to 40% about 1.4 to 2.8 days
1 month not reported not reported
3 months not reported not reported
1 year (365 days) 5% to 10% about 18 to 37 days

Derived by CEOtudent from the two ratios reported by Cepeda and colleagues (2008). Arithmetic: 7 x 0.20 = 1.4 and 7 x 0.40 = 2.8; 365 x 0.05 = 18.25 and 365 x 0.10 = 36.5. The one-month and three-month rows are left unreported because the study anchors only the one-week and one-year points, and interpolating between them would be an invention rather than a finding.

The practical reading: if you need a skill to hold for a week, review it every day or two. If you need it to hold for a year, review it roughly monthly. The proportion collapses by a factor of four while the actual interval grows by more than a factor of ten.

CEOtudent editorial framework: a fast upskilling protocol that respects the constraints

A protocol, not a technique list. Each step names what it is built on and what it is not.

Step What you do Evidence it rests on Honest limit
1. Fix the retention horizon first Decide how long you need the skill to hold before choosing any schedule Cepeda et al. 2008: optimal gap depends jointly on interstudy interval and retention interval Only the one-week and one-year anchors are published
2. Schedule absolute gaps, not a fixed rule Day or two apart for a one-week horizon, roughly monthly for a one-year horizon Derived from the published ratios above Intermediate horizons are not specified by the research
3. Block the vocabulary, interleave the cases Learn terms and definitions in blocks; mix problem types once the terms are known Brunmair and Richter 2019: g = -0.39 favouring blocking for words, g = 0.42 overall interleaving The material categories in the meta-analysis are coarser than real work
4. Replace re-reading with attempted recall Every review session starts with recall from memory before any source is opened Donoghue and Hattie 2021: practice testing d = 0.74 against re-reading d = 0.47; Adesope et al. 2017 found practice tests beat restudying and all other comparison conditions Both were measured mostly on factual recall
5. Force feedback into the loop Check every recall attempt against a source or a person immediately Donoghue and Hattie 2021 name the presence of feedback as an important moderator The paper reports it as a moderator, not as a separate effect size
6. Supply the structure your employer does not Build rotation and peer-review into your own week if the workplace provides neither Eurostat CVTS 2020: job rotation 12.7%, learning circles 13.4%, self-directed learning 29.1% Survey covers EU enterprises; it describes provision, not effectiveness

This is a CEOtudent framework. The evidence column is quoted; the sequencing is editorial judgement.

Step 6 is the one that follows from the deployment data rather than the laboratory data, and it is the one most upskilling advice omits entirely. If the survey says that fewer than one enterprise in seven runs the formats that produce the best-measured conditions, then waiting for your employer to schedule your spacing is waiting for a service that the market does not supply. Our work on how fast abilities fade without practice covers the maintenance side of the same problem, and reskilling after 40 deals with the age-related moderators this piece only touches.

What this does not tell you

The Eurostat survey measures provision, not outcome. It records which formats enterprises offer, not whether anyone learned anything. No part of the deployment data is evidence that job rotation works better than a course inside a real organisation, only that the format structurally contains conditions the laboratory literature favours.

The two datasets are not statistically comparable. Effect sizes come from controlled studies, mostly with students, mostly on factual material. Deployment shares come from a survey of enterprises. Placing them in one table is an analytical move, and the mapping between technique and format is judgement. It is offered as a way of seeing the gap, not as a measurement of it.

The meta-analytic figures carry high heterogeneity. The I-squared values reported in the Donoghue and Hattie table run from 68% to 89% across the ten techniques, which means the averages sit on top of wide variation between studies. A published d of 0.85 is a central tendency across a scattered literature, not a number you should expect to reproduce.

Finally, the 2020 Eurostat reference year covers a period disrupted for most European employers. The 2015 to 2020 movements, particularly the fall in courses and conferences and the rise in self-directed learning, should be read with that in mind. The longer 2005 to 2020 trend for self-directed learning and job rotation is less exposed to that caveat.

FAQ

Which single technique should I use if I only adopt one?
On the measured evidence, distributed practice, at d = 0.85 in the Donoghue and Hattie table, with practice testing second at d = 0.74. The caveat attached to both is the same: they were largely measured on surface and factual outcomes, so treat the ranking as a guide to sequencing rather than a promise about complex skills.

Is interleaving overrated?
It is over-generalised. Brunmair and Richter measured a moderate overall effect at g = 0.42, but also an advantage for blocked practice on word-based material at g = -0.39 and non-significant results for expository texts. The instruction that survives the evidence is conditional: interleave problem types and cases, block vocabulary and definitions.

How often should I review something I need to keep for a year?
The published ratio is 5 to 10% of the retention interval, which works out at roughly 18 to 37 days for a one-year horizon. Monthly review is the practical reading. For a one-week horizon the same study’s ratio of 20 to 40% works out at about one and a half to three days.

Why does my employer’s training not follow any of this?
Because the formats that structurally deliver spacing and retrieval are the ones least commonly provided. In 2020, 54.9% of EU enterprises ran courses while 12.7% ran job rotation and 13.4% ran learning circles. The course is administratively simple and compresses the exposure that the evidence says should be spread out.

Does the research apply to experienced professionals?
Less cleanly than the headline numbers imply. Donoghue and Hattie report that effects were much greater for lower than higher ability learners, and that near versus far transfer was an important moderator. An experienced professional learning an adjacent skill is in the part of the distribution where the measured lift is smaller, which is an argument for tighter protocol design rather than for ignoring the protocol.

Is self-directed learning a good sign or a bad one?
Both. Its rise from 9.9% to 29.1% of EU enterprises between 2005 and 2020 means more autonomy over what you learn and, at the same time, that nobody else is designing the schedule. The techniques with the strongest evidence are precisely the ones that require a schedule.

Sources and further reading

  • Donoghue, G. M. and Hattie, J. A. C. (2021). A Meta-Analysis of Ten Learning Techniques. Frontiers in Education, volume 6, article 581216. Effect sizes, utility classifications, case and participant counts, moderator findings.
  • Brunmair, M. and Richter, T. (2019). Similarity matters: A meta-analysis of interleaved learning and its moderators. Psychological Bulletin, volume 145, issue 11, pages 1029 to 1052. Overall and material-specific interleaving effects.
  • Cepeda, N. J., Vul, E., Rohrer, D., Wixted, J. T. and Pashler, H. (2008). Spacing effects in learning: a temporal ridgeline of optimal retention. Psychological Science, volume 19, issue 11, pages 1095 to 1102. Optimal gap as a proportion of test delay.
  • Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T. and Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, volume 132, issue 3, pages 354 to 380. Scope of the distributed practice literature and the joint operation of interstudy and retention intervals.
  • Adesope, O. O., Trevisan, D. A. and Sundararajan, N. (2017). Rethinking the Use of Tests: A Meta-Analysis of Practice Testing. Review of Educational Research, volume 87, issue 3, pages 659 to 701. Practice testing against restudying and other comparison conditions.
  • Eurostat, Continuing Vocational Training Survey, enterprises providing training by type of training and size class, percentage of all enterprises, EU 27 countries from 2020, reference years 2005, 2010, 2015 and 2020.
  • Rowland, C. A. (2014). The effect of testing versus restudy on retention: a meta-analytic review of the testing effect. Psychological Bulletin, volume 140, issue 6, pages 1432 to 1463. Cited for the finding that initial recall tests yield larger testing benefits than recognition tests.

This content was compiled with the support of AI following in-depth research, then written and prepared for publication by the CEOtudent editorial team.

This post is also available in: Türkçe Français Español Deutsch

Benzer içerikler