The Duolingo Streak: What the Number Is Actually Doing

What a streak counter measures, what the independent evidence supports, and what happens the day it breaks.

Duolingo has published no peer-reviewed research on streaks, and its own A/B tests put the effect at a fraction of a percent. What the independent literature actually shows.

By Human Operating System·September 9, 2026·10 min read
Stack of identical glass plates with one fractured plate slipping out of alignment under gold light

The streak is the most effective thing Duolingo ever built, and it teaches you nothing.

That is not a criticism of the app. It is a description of what a streak counter is: a running total of consecutive days on which you opened it. It measures attendance. Whether anything was learned on any of those days is a separate question that the number does not track and cannot answer.

The interesting question is not whether the streak works - it plainly does something, or the company would not have built its entire retention strategy around it. The interesting question is what it works on, how strong that mechanism actually is once you look at the independent literature, and what it does to you on the day it breaks.

What Duolingo has actually published

Start here, because it changes how you read everything else.

Duolingo has published no peer-reviewed research on streaks. Its research index lists papers at ACL, EMNLP, KDD, EDM, NAACL-HLT, CogSci, Psychological Science, Language Learning and Cognitive Science. They are about spaced repetition, adaptive testing, psychometrics, notification scheduling and natural language processing. Not one is about streaks, streak freezes, or habit formation.

Every quantitative claim the company makes about streaks lives on its marketing blog. No data released, no confidence intervals, no p-values, no independent audit.

That does not make the claims false. It does mean they should be read as what they are. And when you read them closely, something instructive appears - the correlational numbers are enormous and the causal numbers are tiny.

The correlational claims (2022 and undated blog posts): learners who reach a seven-day streak are "3.6 times more likely to complete their course"; elsewhere, "2.4 times more likely to continue using Duolingo the next day." Note these are different outcomes with different multipliers, so neither can be restated as the other. Both are raw correlations with an obvious problem: people who make it to day seven were already more motivated than people who did not.

The causal claims - the company's own A/B tests, which are the numbers that actually isolate the effect of a design change - come out at +0.38% daily active learners from doubling streak freezes, +1.7% seven-day retention from adding streak animations, +1% overall daily active learners and +3.3% day-14 retention from a streak redesign.

One to two orders of magnitude smaller. That gap between the correlational story and the experimental result is the most honest thing in Duolingo's public record on streaks, and it is not in any of the headlines.

Loss aversion is the standard explanation, and it is the weakest link

Ask anyone why streaks work and you will get the same answer: loss aversion. You have 200 days; losing them would hurt more than gaining day 201 would please you; so you open the app.

The problem is that loss aversion is no longer the settled fact it is usually presented as.

The lambda = 2.25 figure everyone quotes ("losses hurt about twice as much as gains feel good") comes from Tversky and Kahneman's 1992 cumulative prospect theory paper, not from the famous 1979 Econometrica paper it is usually attributed to. And the last few years have produced a genuine fight over whether it holds:

  • Brown, Imai, Vieider and Camerer (2024), in the Journal of Economic Literature, meta-analysed 607 estimates from 150 articles and found a mean coefficient of 1.955, 95% CI [1.820, 2.102]. This supports loss aversion.
  • Walasek, Mullett and Stewart (2024), in the Journal of Economic Psychology, meta-analysed 19 datasets from 17 articles in risky choice and found lambda = 1.31, 95% CI [1.10, 1.53]. They also reported that "the confidence intervals incorporated loss neutrality in 12 of the 19 datasets," and that "much of the data are of poor quality."
  • Yechiam and Zeif (2025) re-analysed Brown et al.'s own dataset, splitting studies by design features. In studies with symmetric gains and losses and no ordering of items - the cleaner designs - the coefficient was approximately 1.07 and not significantly above 1.0.

There is also a specific finding that bears directly on streaks. Yechiam (2019) concluded that "while very large losses are overweighted, smaller losses are often not." A lost streak is a small, symbolic, non-monetary loss. It is precisely the category the challenge to loss aversion is strongest about.

And there is a neat piece of evidence from inside the loyalty-card literature itself. Kivetz, Urminsky and Zheng (2006) tested loss aversion as the mechanism behind illusory goal progress and found it did not explain the effect - the difference between their experimental and control conditions on a loss-aversion emotional scale was t = .2, not significant.

So: loss aversion may be part of the story. It is not the confident, quantified explanation it is usually offered as, and anyone telling you the streak works because losses loom twice as large is citing a number from 1992 that the 2020s have not been kind to.

The goal gradient: real, partial, and resting on 108 people

The better-supported mechanism is the goal gradient - people accelerate as they approach a reward.

Kivetz, Urminsky and Zheng's cafe study is the canonical demonstration. Tracking 949 completed ten-stamp coffee cards across roughly 10,000 purchases, they found customers bought more frequently as they neared a free coffee: the gap between first and last inter-purchase time was 0.7 days, "representing an average acceleration of 20%" (t = 2.6, p < .05).

The famous part is the second study. 108 customers were given either a ten-stamp card or a twelve-stamp card that arrived with two stamps already filled in. Both required ten actual purchases. The pre-stamped group finished in 12.7 days against 15.6 - "nearly three days or 20% faster (t = 2.0, p < .05)."

This is one of the most-cited results in consumer psychology and it is worth knowing its dimensions: n = 108, t = 2.0, barely under p < .05, one field experiment, 2006. We could not find any published direct replication, successful or failed.

The larger study also contains a detail the retellings drop. A latent-class model found the acceleration in one class of customers (58%) and not reliably in the others: one class was non-significant at p = .13, and two more had a goal-distance coefficient near zero. Roughly 42% of customers showed a weak or absent goal gradient. The effect is real. It is not universal, and it may well not be yours.

Gamification: the part that survives methodological scrutiny is not the motivating part

Three meta-analyses put gamification's effect on learning outcomes at around g = 0.46-0.50: Sailer and Homner (2020) at g = .49 for cognitive outcomes; Bai, Hew and Huang (2020) at g = 0.504, 95% CI [0.284, 0.723] across 24 quantitative studies and 3,202 participants; Huang et al. (2020) at g = .464, 95% CI [.244, .684] across 30 studies and 3,083 participants. That is genuine convergence.

But Sailer and Homner did something the others did not: they re-ran the analysis on only the methodologically rigorous studies. The result:

OutcomeAll studiesHigh methodological rigour only
Cognitiveg = .49, 95% CI [0.30, 0.69]g = .42, 95% CI [0.14, 0.68] - still significant
Motivationalg = .36, 95% CI [0.18, 0.54]g = .22, p = .20 - not significant
Behaviouralg = .25, 95% CI [0.04, 0.46]g = .27, p = .22 - not significant

The learning effect holds. The motivational and behavioural effects do not survive the quality filter. A streak counter is a motivational and behavioural device, not a cognitive one, so this is the subsplit that matters most here - and it is the one that fails.

There is a further limit worth stating. Sailer and Homner note that only one study in their entire meta-analysis reported a delayed post-test. The literature is essentially silent on whether any of these effects persist beyond the novelty period. That is not "novelty effects have been ruled out" and it is not "novelty explains it all." It is that nobody has looked.

What happens when it breaks

This is the part with the strongest independent evidence, and it is the part Duolingo's blog does not discuss.

Silverman and Barasch (2023), in the Journal of Consumer Research, ran seven studies on logged streaks. Their central finding is that an intact streak in a behaviour log increases subsequent engagement - in a fitness-app field dataset of 980 users (b = 0.38, SE = .03, Z = 11.27, p < .001), and in controlled studies at 66.23% versus 57.86% (chi-squared(1) = 4.47, p = .035) and 65.84% versus 47.98% (chi-squared(1) = 13.02, p < .001).

Their third study is the one to read carefully. It compared people with a broken streak who were shown the log against people with the identical behavioural history who were shown no log at all:

participants with a broken streak were less likely to engage in the target behavior when their broken streak was highlighted via the log (45.21%) compared to when it was not (60.90%; chi-squared(1) = 7.46, p = .006, OR = 0.53)

Being shown your broken streak left people worse off than not being tracked at all. Same behaviour, same history, different display, fifteen percentage points of engagement.

The authors are explicit that the effect "is independent of actual past behavior and depends solely on how that behavior is represented within the log." The streak is not measuring your commitment. It is manufacturing a goal out of a representation, and when the representation breaks, the manufactured goal breaks with it.

Report the contrary evidence too. A randomised controlled trial by Aulagnon, Cristia, Cueto and Malamud (2024), run on a Peruvian maths platform with 60,000 students randomised, found highlighting streaks raised the likelihood of connecting by 2.8 percentage points and stated: "We do not observe a discouragement effect from highlighting streaks." That cuts against Silverman and Barasch. Two caveats: the trial was not designed to isolate what happens at the moment a streak breaks, its achievement estimates rest on the 1,503 students of 60,000 who completed an endline test, and it is a working paper, not peer-reviewed.

The streak freeze is the best-supported part of the whole system

Here is the irony. The independent literature supports Duolingo's forgiveness mechanic more clearly than it supports the streak counter itself.

Sharif and Shu (2017), in the Journal of Marketing Research, studied what they called "emergency reserves" - an explicitly defined allowance of slack within a goal. Across six studies they found reserves were both preferred and associated with greater persistence. The mechanism is the elegant part: people persist "because they want to avoid using the 'emergency' reserve."

And Silverman and Barasch found the damage from a broken streak was "attenuated when consumers can 'repair' a broken streak."

So the peer-reviewed evidence says: a mechanism that lets you miss a day without losing everything is good for persistence, and it works partly by being something you would rather not spend. That is the streak freeze, described by independent researchers who were not studying Duolingo.

Does the streak mean you are learning Spanish?

No, and it is worth separating three claims that get collapsed into one.

Does Duolingo teach language? The evidence is thinner than the marketing suggests, and nearly all of the big numbers are company-funded.

The famous "34 hours equals a college semester" figure comes from Vesselinov and Grego (2012), a study Duolingo funded, never peer-reviewed, with a final sample of 88 people - average age 34.9, and 69.3% college graduates or higher. Its own report gives a mean gain of 8.1 WebCAPE points per hour against a median of 3.9. The distribution is heavily skewed; the typical participant gained less than half the headline rate. Krashen (2014), critiquing it independently, noted the sample was unrepresentative and the outcome measure "a multiple-choice test that is clearly form-based."

The "equivalent to four semesters" framing traces to Jiang, Rollinson, Plonsky, Gustafson and Pajak (2021) in Foreign Language Annals, where four of five authors were Duolingo employees, Duolingo paid for the tests and the participant compensation, and the design was posttest-only with no control group. The abstract states plainly: "No other skills were assessed" - no speaking, no writing.

The one study we found with genuine comparison groups is Kim, Payant, Skalicky and Namkung (2026) in Studies in Second Language Acquisition: 183 learners across classroom-only, Duolingo-only and combined conditions over 16 weeks. Its conclusion is far more modest and far more credible: "both traditional classroom instruction and Duolingo are comparably effective for beginning French language learners." Comparable to a class. Not better than one, and not a substitute for the skills it does not test.

Does the streak make you learn more?

Nothing shows this. There is one adjacent finding worth reporting fairly: a Duolingo-distributed whitepaper by Plonsky and Sudina (2023), studying 287 learners over six months, found that frequency of engagement correlated with gains more reliably than total minutes, and advised learners to "log in to the app frequently... even if those sessions do not last for long periods of time." That is the closest thing to evidence for the streak's underlying logic. It is correlational, it is distributed by the company, and it says nothing about whether a counter causes the frequency.

Does the streak make you open the app?

Yes. That much is not in doubt, and it is what the streak was built to do.

So should you keep it?

Two honest positions, and the evidence supports both.

Keep it

if the streak is getting you to sit down more days than you otherwise would, and if you can lose it without the loss taking the habit with it. Frequency does seem to matter more than session length, and a device that produces frequency is doing something real. Use the streak freeze without guilt - it is the best-evidenced element in the whole system, and its whole design is to let you miss a day without the display telling you that you failed.

Let it go

if you have noticed the shape Silverman and Barasch measured: doing the minimum lesson to keep the number rather than the lesson you would learn from, or the flat, deflated feeling after a break that makes it harder to come back than if you had never counted. That is not a failure of discipline. It is the documented behaviour of a system where "maintaining a logged streak" has quietly become a goal in its own right, separate from the one you started with.

The number is not measuring you. It is measuring itself.

For the wider question of how interfaces are built to hold attention, see the persuasive design guide.
For whether study apps in general earn their place, see study tools and apps actually worth it.

Recommended Resource
#duolingo streak#streaks#gamification#loss aversion#goal gradient#attention design
About the Author

Human Operating System

Human Operating System is a research-led publication about the human mind under digital pressure. We report what the evidence does - and does not - support.

About Human Operating System

Sources & Further Reading

17 sources

These are the sources used for this article. Where a study's limits matter to the claim, those limits are kept in the citation.

View all 17 sourcesHide sources
  1. Silverman, J., & Barasch, A. (2022). On or Off Track: How (Broken) Streaks Affect Consumer Decisions. Journal of Consumer Research, 49(6), 1095-1117. Open source ↗
  2. Kivetz, R., Urminsky, O., & Zheng, Y. (2006). The Goal-Gradient Hypothesis Resurrected: Purchase Acceleration, Illusionary Goal Progress, and Customer Retention. Journal of Marketing Research, 43(1), 39-58. Open source ↗
  3. Sharif, M. A., & Shu, S. B. (2017). The Benefits of Emergency Reserves: Greater Preference and Persistence for Goals that Have Slack with a Cost. Journal of Marketing Research, 54(3), 495-509. Open source ↗
  4. Sailer, M., & Homner, L. (2020). The Gamification of Learning: a Meta-analysis. Educational Psychology Review, 32, 77-112. Open source ↗
  5. Bai, S., Hew, K. F., & Huang, B. (2020). Does gamification improve student learning outcome? Educational Research Review, 30, 100322. Open source ↗
  6. Huang, R., Ritzhaupt, A. D., Sommer, M., Zhu, J., Stephen, A., Valle, N., Hampton, J., & Li, J. (2020). The impact of gamification in educational settings on student learning outcomes: a meta-analysis. Educational Technology Research and Development, 68, 1875-1901. Open source ↗
  7. Brown, A. L., Imai, T., Vieider, F. M., & Camerer, C. F. (2024). Meta-analysis of Empirical Estimates of Loss Aversion. Journal of Economic Literature, 62(2), 485-516. Open source ↗
  8. Walasek, L., Mullett, T. L., & Stewart, N. (2024). A meta-analysis of loss aversion in risky contexts. Journal of Economic Psychology, 103, 102740. Open source ↗
  9. Yechiam, E. (2019). Acceptable losses: the debatable origins of loss aversion. Psychological Research, 83, 1327-1339. Open source ↗
  10. Yechiam, E., & Zeif, D. (2025). Loss aversion is not robust: A re-meta-analysis. Journal of Economic Psychology, 107, 102801.
  11. Gal, D., & Rucker, D. D. (2018). The Loss of Loss Aversion: Will It Loom Larger Than Its Gain? Journal of Consumer Psychology. Open source ↗
  12. Kim, Y., Payant, C., Skalicky, S., & Namkung, Y. (2026). Comparing the effectiveness of Duolingo, Classroom instruction, and Classroom + Duolingo instruction conditions on beginner-level French language development. Studies in Second Language Acquisition, First View. Open source ↗
  13. Jiang, X., Rollinson, J., Plonsky, L., Gustafson, E., & Pajak, B. (2021). Evaluating the reading and listening outcomes of beginning-level Duolingo courses. Foreign Language Annals, 54(4), 974-1002. Open source ↗
  14. Vesselinov, R., & Grego, J. (2012). Duolingo Effectiveness Study: Final Report. (Funded by Duolingo; not peer-reviewed.)
  15. Krashen, S. (2014). Does Duolingo Trump University-Level Language Learning? The International Journal of Foreign Language Teaching.
  16. Aulagnon, R., Cristia, J., Cueto, S., & Malamud, O. (2024). Streaking to Success: The Effects of Highlighting Streaks on Student Effort and Achievement. IDB Working Paper No. IDB-WP-1566. (Working paper; not peer-reviewed.)
  17. Plonsky, L., & Sudina, E. (2023). The effects of frequency, duration, and intensity on L2 learning through Duolingo. Duolingo-hosted whitepaper.
The Weekly System

The Weekly System

Join the launch list for one calm, research-led email about attention, memory, learning and the systems designed to hold your attention. No noise, no panic and no unsupported certainty.