Spaced Repetition: The Effect Is Real, the Rules People Quote Are Not
Spacing study sessions is one of the best-supported findings in learning research, but the popular rules about expanding intervals and fixed review ratios do not survive the evidence.

Spacing your study sessions instead of cramming them together is about as close to settled as educational psychology gets. It has been demonstrated since the 1880s, replicated in hundreds of experiments, and it survives meta-analysis comfortably.
What has not survived is most of what gets said about how to space. The two rules you will meet most often - that expanding intervals beat fixed ones, and that you should review at a fixed percentage of your target retention interval - are respectively a well-powered null and a misreading of the very study it comes from.
The effect itself
Cepeda, Pashler, Vul, Wixted and Rohrer (2006) synthesised 317 experiments from 184 articles, covering work from 1885 to 2002, with 839 assessments of distributed practice.
Spacing beat massing in 259 of 271 comparisons. Only twelve showed no benefit. At retention intervals under a minute the improvement was around 9%; for studies testing after more than a month, the average benefit rose to about 15%.
The authors flag their own limitation clearly, and it matters: roughly 80% of comparisons used retention intervals under one day, and only about 4% exceeded a month. The effect is best evidenced at exactly the timescale students care about least.
A more recent synthesis by Donoghue and Hattie (2021) - 242 studies, 1,619 effects, 169,179 participants - puts distributed practice at d = 0.85, among the highest-ranked strategies alongside practice testing at d = 0.74. With the same caveat attached: outcomes were predominantly factual, and 74% of studies completed in a day or less.
The number everyone quotes, and what it actually says
Cepeda and colleagues (2008) ran a large randomised internet experiment - 1,354 participants completing all three sessions - across 26 combinations of study gap and retention interval.
The headline is genuinely impressive: the optimal gap produced a 64% increase in final recall and a 26% increase in recognition, compared with studying twice in immediate succession.
But the part that gets dropped is the shape. The optimal gap, expressed as a proportion of the retention interval, was not constant:
| Retention interval | Optimal gap | Gap as % of interval |
|---|---|---|
| 7 days | ~1 day | ~14% |
| 35 days | ~11 days | ~31% |
| 70 days | ~21 days | ~30% |
| 350 days | ~21 days | ~7% |
The ratio rises, then falls sharply. There is no fixed percentage. The widely repeated advice to “review at 10-20% of your target interval” takes a non-linear function and reports one point on it as a rule.
The practical implication is more useful than the rule anyway: the absolute optimal gap barely moves between 70 days and 350 days. If you want to remember something for a year rather than ten weeks, the schedule that gets you there is nearly the same one. And the function is flat near its peak, so being roughly right is nearly as good as being exactly right - which is why the precision claimed by scheduling algorithms is mostly decorative.
The expanding-schedule myth
Most flashcard software expands its intervals: review after a day, then three days, then a week, then a month. The rationale is intuitive and the evidence for it is a null.
Latimier, Peyre and Ramus (2021) meta-analysed both questions separately.
Spaced versus massed retrieval practice - 39 effect sizes from 11 studies - gave g = 1.01, 95% CI [0.68, 1.34]. Publication bias was detected (Egger’s t = 4.41, p < .0001), and trim-and-fill reduced it to g = 0.74 [0.55, 0.91]. Still large.
Expanding versus uniform schedules - 54 effect sizes from 16 studies - gave g = 0.034, 95% CI [-0.10, 0.17], p = .59. Heterogeneity I2 = 0%, no publication bias detected.
That is a precise, well-powered, homogeneous null. Expanding intervals are not better than fixed ones.
The primary study behind the confusion is instructive. Karpicke and Roediger (2007) found expanding schedules beat equal ones at a 10-minute delay - 71% against 62%. At two days the result reversed: 33% against 45%. Their third experiment isolated the active ingredient, and it was not the expansion: it was delaying the first test.
So the thing that works is putting a gap before your first review. What happens after that appears not to matter much.
Do the apps work?
This is where the evidence thins, and it is worth being blunt about it.
There is no randomised controlled trial of Anki, Quizlet, Duolingo or SuperMemo against a control condition that we were able to verify. The evidence people cite for these products is of three weaker kinds.
Observational analysis of app data. Tabibian and colleagues (2019) analysed roughly 5.2 million user-word pairs from Duolingo and reported that learners whose schedules more closely resembled an optimal algorithm memorised more effectively. But users chose their own schedules - this is a natural experiment, not a trial, and self-selection is unexcluded. We could not retrieve the results section, so we are printing no percentage improvement from it.
Cohort comparisons. Gilbert and colleagues (2023) compared 78 Anki users against 52 non-users among first-year medical students and found MCAT-adjusted advantages of 6.2 to 12.9 percentage points. The authors concede the obvious: no randomisation, and “more motivated students using additional resources.”
One genuine randomised trial - of a different app. Upadhyay and colleagues (2021) randomised 11,481 learners inside a German driving-theory app across machine-learning, difficulty-ordered and random scheduling, covering about 16.75 million answers. The algorithmic condition produced roughly a 48% reduction in forgetting rate against random scheduling and a 92% longer median half-life. Their own caveat is important: the algorithmic group was more likely to quit within two days. An efficiency gain paid for with adherence is not obviously a gain.
The classroom problem
Almost all of the strong evidence is laboratory work on word lists. When spacing is taken into real courses, it shrinks.
A 2024 within-subjects study across nine introductory STEM courses with 578 undergraduates compared spaced against massed retrieval practice over six weeks. The eight-course pooled effect was +1.50%, non-significant. Adding the ninth course gave +2.06%, 95% CI [0.16, 3.97] - statistically significant and educationally trivial. Heterogeneity was I2 = 89.2%, and the effect reached significance in only two of nine courses.
That is the honest counterweight to a 64% laboratory gain, and any article quoting the 64% without it is misleading you.
A smaller web-application study (Belardi et al., 2021, N = 79) found spacing worth a large jump in vocabulary recall - four sessions 77.2% against one session 52.5%, F(2,76) = 8.51, p = .0005 - but assignment to spacing conditions was not random, for scheduling reasons. Its two nulls are worth having: testing (p = .58) and multimodality (p = .61) did nothing in that design.
What this does not establish
That spacing transfers reliably to complex material. The lab evidence is dominated by vocabulary and factual recall. The classroom evidence is weak and heterogeneous.
That any particular app is effective. No product in this category has a controlled trial behind it that we could verify.
A schedule you should follow. The optimal-gap function is non-linear and flat near its peak; there is no ratio to apply.
That expanding intervals help. They meta-analyse to zero.
What actually follows
Put a gap before your first review. This is the piece with the strongest and most consistent support - Karpicke and Roediger’s third experiment isolated it, and the whole distributed-practice literature rests on it.
Do not agonise over the schedule. The function is flat near its optimum, expanding beats fixed by nothing, and the difference between a good schedule and a perfect one is smaller than the difference between reviewing and not reviewing.
Space over weeks, not hours, if you want to keep it. Cepeda’s data show the absolute optimal gap for a one-year horizon is roughly three weeks - and 80% of the underlying literature never tested past a day, so the long-horizon evidence is thinner than the effect’s reputation suggests.
Expect classroom gains an order of magnitude smaller than lab gains. A 64% improvement on word pairs became about 2% across nine real courses. Both numbers are real; only one of them describes studying.
Spacing is one of the few study techniques that survives testing - see the cognitive offloading guide for the wider picture, sans forgetica for one that did not, and study tools and apps actually worth it for the software question.
Human Operating System
Human Operating System is a research-led publication about the human mind under digital pressure. We report what the evidence does - and does not - support.
About Human Operating SystemSources & Further Reading
10 sourcesThese are the sources used for this article. Where a study's limits matter to the claim, those limits are kept in the citation.
View all 10 sourcesHide sources
- Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132(3), 354–380. Open source ↗
- Cepeda, N. J., Vul, E., Rohrer, D., Wixted, J. T., & Pashler, H. (2008). Spacing effects in learning: A temporal ridgeline of optimal retention. Psychological Science, 19(11), 1095–1102. Open source ↗
- Latimier, A., Peyre, H., & Ramus, F. (2021). A meta-analytic review of the benefit of spacing out retrieval practice episodes on retention. Educational Psychology Review, 33, 959–987. Open source ↗
- Karpicke, J. D., & Roediger, H. L. (2007). Expanding retrieval practice promotes short-term retention, but equally spaced retrieval enhances long-term retention. Journal of Experimental Psychology: Learning, Memory, and Cognition, 33(4), 704–719. Open source ↗
- Upadhyay, U., Lancashire, G., Moser, C., & Gomez-Rodriguez, M. (2021). Large-scale randomized experiments reveals that machine learning-based instruction helps people memorize more effectively. npj Science of Learning, 6, 26. Open source ↗
- Tabibian, B., Upadhyay, U., De, A., Zarezade, A., Schölkopf, B., & Gomez-Rodriguez, M. (2019). Enhancing human learning via spaced repetition optimization. PNAS, 116(10), 3988–3993. Open source ↗
- Gilbert, M. M., et al. (2023). A cohort study assessing the impact of Anki as a spaced repetition tool on academic performance in medical school. Medical Science Educator, 33, 955–962. Open source ↗
- Belardi, A., Pedrett, S., Rothen, N., & Reber, T. P. (2021). Spacing, feedback, and testing boost vocabulary learning in a web application. Frontiers in Psychology, 12, 757262. Open source ↗
- Donoghue, G. M., & Hattie, J. A. C. (2021). A meta-analysis of ten learning techniques. Frontiers in Education, 6, 581216. Open source ↗
- Single-paper meta-analyses of spaced retrieval practice in nine introductory STEM courses (2024). International Journal of STEM Education, 11. Open source ↗
The Weekly System
Join the launch list for one calm, research-led email about attention, memory, learning and the systems designed to hold your attention. No noise, no panic and no unsupported certainty.
