TL;DR: By design, Sesame Street was less an entertainment programme than a learning experiment. The production team and education researchers worked inside the same process, episodes were measured on children before they aired, and the results changed the next episode. That loop was later named the CTW model, and the programme’s first two seasons were separately evaluated by an independent body, the Educational Testing Service. The strongest evidence for the experiment, however, arrived fifty years later from an unexpected direction. A study by Melissa Kearney and Phillip Levine, published in the American Economic Journal: Applied Economics in January 2019, used a technical accident that made the 1969 broadcast receivable in some areas and not others, the difference between UHF and VHF transmission, as a natural experiment. In areas where the signal reached, preschool-age children were 16 percent less likely to fall behind grade level in elementary and middle school. The effect was strongest among boys and among children in economically disadvantaged areas. The same study states plainly that its estimates for long-term educational and labour market outcomes are imprecise. Below: the design of the experiment, a table of verified findings, and an original map of lessons for anyone building their own learning system.
It seems odd for a children’s programme to be among the most heavily researched productions ever made. For Sesame Street it is not odd. It was the design.
When the programme was being prepared in the late 1960s, the team was deliberately split in two. On one side were experienced television producers. On the other were sociologists, pedagogues, psychologists and researchers who studied learning. Having those two groups work inside the same production process was unusual for the period, and it was the programme’s real innovation.
The purpose behind the project was not entertainment either. Children from families without economic means could not access preschool education and were therefore starting school behind. Sesame Street was an attempt to close that gap with a broadcast signal.
The real design: a loop that measured and changed
The way of working that emerged by the end of the first season was later called the CTW model. It combined planning, production and evaluation in a single loop: researchers and early childhood educators worked in the same process as the writers and directors.
That model had two legs, and it matters not to confuse them.
The first was in-house formative research. Its purpose was to make the episode better. Researchers invented their own tools to measure how much attention young viewers gave the screen. If a scene did not hold attention, the scene changed.
The second was independent summative evaluation. Its purpose was to judge from outside whether the programme worked at all. That job went to the Educational Testing Service for the first two seasons. The first-year evaluation prepared by Samuel Ball and Gerry Ann Bogatz was published in 1970; the second followed in 1971.
That distinction is directly useful to anyone designing their own learning. Feedback that improves the material and measurement that judges whether the material worked are not the same thing, and both degrade when the same person does both.
The evidence that arrived fifty years later
The early evaluations were positive but could not answer one question: might the children who watched simply have come from different families? Was the observed difference caused by the programme, or by the characteristics of the families who tuned in?
In 2019 two economists resolved that question using an accident nobody had planned.
In 1969, Sesame Street could not be received at the same quality everywhere. In some areas the broadcast came over VHF; in others it came over UHF, which the televisions of the period received far worse. Signal quality varied by geography and technology, not by a family’s income, education or interest. That was exactly the kind of natural experiment researchers want.
Table 1. Verified research findings on Sesame Street
| Finding | Verified content | Source |
|---|---|---|
| Method | Geographic variation in reception arising from UHF versus VHF transmission, matched to Census data | Kearney and Levine, AEJ: Applied Economics, volume 11, issue 1, January 2019, pages 318-350 |
| Main result | School performance improved where the programme could be received, particularly for boys | Same study |
| Size of the effect | A 16 percent reduction in the likelihood that children would fall behind in elementary and middle school | Kearney, AEA research interview |
| Honest caveat | Point estimates for long-term educational and labour market outcomes are generally imprecise | Kearney and Levine, article abstract |
| Reach | In any given week, half of two- to five-year-olds were watching the show | Kearney, AEA research interview |
| Reach in disadvantaged areas | Nearly 90 percent of children were watching | Kearney, AEA research interview |
| Comparison | Early evaluations found literacy and numeracy benefits on par with Head Start, though Head Start also provided comprehensive services beyond educational content | Kearney, AEA research interview |
Do not skip the fourth row. The study itself says its estimates for long-term outcomes are imprecise. Sesame Street improved school readiness; a sentence like “children who watched went on to earn more years later” does not come out of this data. Not loading a finding with more than it can carry matters as much as the finding itself.
The reach figures are striking in their own right. We are talking about a programme watched by nearly 90 percent of children in disadvantaged neighbourhoods. An educational intervention with reach that wide produces a large aggregate difference even with a modest individual effect. It is one of the cleanest illustrations available of the relationship between scale and effect size.
What this gives someone designing their own learning today
The table below is this piece’s original contribution. On the left is the design principle from the experiment, on the right its translation to individual learning and the typical error that appears when the principle is violated.
Table 2. From the Sesame Street experiment to principles of individual learning (CEOtudent editorial framework)
| # | Principle from the experiment | Individual equivalent | Error when violated |
|---|---|---|---|
| 1 | Producers and researchers worked in the same process | The learning plan and the measure of success are defined at the same moment | Discovering what you learned much later, often never |
| 2 | Formative research was separated from summative evaluation | Feedback that improves the material is kept apart from the test of whether it worked | Mistaking practice for a test and overrating your own performance |
| 3 | Attention was measured and scenes changed accordingly | Record where concentration breaks, and rebuild the working method around it | Treating lost focus as a willpower problem and repeating the same method |
| 4 | Evaluation was contracted to an outside body | What has been learned is tested against an external, independent standard | Grading yourself and finding yourself sufficient |
| 5 | Reach came before effect | Establish a sustainable repetition frequency before hunting for the perfect method | Never starting while trying to build the flawless system |
| 6 | The target audience was those who needed it most | Effort is directed at the weakest link, not the most enjoyable subject | The false sense of progress from restudying what you already know |
| 7 | Results kept being measured for fifty years | Results are re-checked at long intervals, not in a single test | Mistaking short-term recall for durable learning |
Rows two and four should be read together. The most common error in learning is not failing to test yourself, but testing yourself with the wrong thing. Rereading a text and finding it familiar, then treating that familiarity as knowledge, is exactly the substitution of formative feedback for summative evaluation. The Sesame Street team separated the two institutionally; an individual can separate them by calendar.
Row five draws the most resistance. Sesame Street’s effect came less from the perfection of individual episodes than from reaching nearly every child. The individual equivalent: a mediocre session done five times a week beats a flawless session done once a month almost every time.
The CEO and student read
The CEO half is measurement discipline. What made Sesame Street distinctive was not the content but the fact that the content was measured and the measurement changed the production. For someone running themselves like a company, the question is not “what am I learning” but “how do I measure what I learned, and what does the measurement change”. A system that does not measure is a well-intentioned habit; a system that measures is a system that improves.
The student half is honesty. The most important line in this piece is not the most impressive one: the study itself says the long-term estimates are imprecise. Not loading evidence with more than it can carry is the hardest and most valuable part of learning. If some questions remain open even in a programme measured for fifty years, it is worth reconsidering how much your three-week impression of your own progress can carry.
Frequently asked questions
Was Sesame Street really an experiment?
Not a randomised experiment in the classic sense. But the production process rested on continuous measurement, and the 2019 study used the geographic distribution of the broadcast signal as a natural experiment. That is a design constructed after the fact, but methodologically strong.
What exactly does the 16 percent reduction refer to?
It refers to the reduction in the likelihood that preschool-age children in reception areas would fall behind the grade appropriate for their age in elementary and middle school. It is not a grade average or an intelligence measure; it is an indicator of falling behind at the grade level.
Did the programme raise income or educational attainment in the long run?
The study cannot show this clearly, and says so itself. Point estimates for long-term educational and labour market outcomes are imprecise. An effect on school readiness and an effect on lifetime outcomes are different claims.
Does this contradict current screen-time advice?
The finding here concerns a specific kind of educational content, designed through measurement with explicit learning objectives and age-appropriate structure, not screen time in general. Generalising it to all screen use draws a conclusion the data does not carry.
What is the most practical individual lesson from this experiment?
Row two of Table 2: separate your practice from your test. If the material you learn with and the standard you test yourself against are the same, what you are measuring is familiarity, not learning.
Sources
Melissa S. Kearney and Phillip B. Levine, Early Childhood Education by Television: Lessons from Sesame Street, American Economic Journal: Applied Economics, volume 11, issue 1, January 2019, pages 318-350.
American Economic Association, Learning from Sesame Street, research interview with Melissa Kearney.
Samuel Ball and Gerry Ann Bogatz, The First Year of Sesame Street: An Evaluation, Educational Testing Service program report, 1970.
Gerry Ann Bogatz and Samuel Ball, The Second Year of Sesame Street: A Continuing Evaluation, Educational Testing Service, 1971.
Children’s Television Workshop, institutional documentation on formative research and the production model, 1970s.
This content was compiled with the support of AI following in-depth research, then written and prepared for publication by the CEOtudent editorial team.
This post is also available in:
















