Binge Watching: Nineteen Definitions, One Experiment, and It Was About a Setting

Researchers use nineteen definitions of binge-watching. The only randomised experiment changed a default setting, not the viewer, and daily viewing fell.

By Human Operating System·October 10, 2026·8 min read
A viewer on a sofa in a dark room facing a screen of endlessly receding frames, with a small gold switch on a side table

Almost everything written about binge-watching treats it as a property of the viewer - a habit, a compulsion, a thing you do too much of and should do less of.

The research points somewhere else, twice over.

First, the field cannot agree on what the behaviour is. A systematic review found that the studies which bothered to define binge-watching produced nineteen distinct definitions between them. That is not a footnote about methodology. It means the prevalence figures, the correlations and the risk factors are not all measuring the same thing.

Second, the only randomised experiment anyone has run on it did not change anything about the viewers. It changed a default setting, and viewing time fell.

The definitional problem, stated exactly

Flayelle, Maurage, Ridell Di Lorenzo, Vögele, Gainsbury and Billieux (2020) published the first systematic review of binge-watching in Current Addiction Reports. It covers 24 studies and 17,545 participants.

Twenty-two of the 24 studies directly operationalised binge-watching. Between them they listed 28 definitions, which the reviewers reduced to 19 distinct possibilities.

Nineteen. Some studies counted episodes in a sitting, some counted hours, some counted sessions per week, some asked people whether they considered themselves binge-watchers. Six papers offered two definitions each.

This matters immediately, because of the number everyone quotes.

The 72% figure, and why we are not printing it the way you have seen it

Twelve studies in the review reported a prevalence rate. Those rates ranged from 44.6% to 98%. The review computes their average as 72.14%, and writes that this suggests binge-watching “is not an atypical viewing practice, but rather the norm across the current samples.”

Read what that number actually is. It is an unweighted arithmetic mean of twelve study-level percentages, drawn from convenience samples, where each study used its own definition of the behaviour it was measuring. It is not a population prevalence. It is not a pooled meta-analytic estimate. It is not weighted by sample size.

“72% of people binge-watch” is a misreading, and a common one. The honest reading is stranger and more interesting: the same review that documents nineteen incompatible definitions then averages the prevalences those definitions produced. The number is an artefact of the disagreement, not a measurement of the behaviour.

The field explicitly does not want this pathologised

This is worth stating because it cuts against the tone of most coverage.

The review’s own conclusion is that “future research should maintain the distinction between high and problematic involvement in binge-watching to avoid overpathologizing this common behavior,” and that “high but healthy engagement in TV series watching should be distinguished from problematic binge-watching to avoid pathologizing this highly popular activity.”

Binge-watching appears in neither DSM-5-TR nor ICD-11 in any form. Gambling disorder is the only behavioural addiction recognised as a clinical disorder in DSM-5, with internet gaming disorder listed as requiring further research; ICD-11 categorises gambling and gaming disorders as disorders due to addictive behaviours. Television viewing is not among them.

What the correlational evidence shows, and what it cannot

Alimoradi and colleagues (2022) meta-analysed binge-watching and mental health across 16 studies and 8,077 adults. The associations, as Fisher’s z:

OutcomeAssociationHeterogeneity
Stress.32I2 = 86.9%
Anxiety.25I2 = 89.7%
Loneliness.19I2 = 65.2%
Insomnia.16I2 = 88.4%
Depression.14I2 = 90.6%

Two facts about that table govern how it can be read.

Every one of the sixteen studies is cross-sectional. The authors state this precludes causal inference. Nothing here establishes that binge-watching causes stress rather than stressed people watching more television.

The heterogeneity is severe - between 65% and 91%. The authors flag that the use of different instruments “carries a risk of measuring different constructs,” which is the definitional problem arriving in the statistics.

Exelmans and Van den Bulck (2017), in the Journal of Clinical Sleep Medicine, surveyed 423 young adults aged 18 to 25, of whom 80.6% identified as binge viewers. Binge viewing was associated with poorer sleep quality (β = .145, p < .01), more daytime fatigue (β = .131, p < .05) and more insomnia symptoms (β = .161, p < .01), with binge viewers having a 98% higher likelihood of poor sleep quality (Exp(B) = 1.981, p < .05). Cognitive pre-sleep arousal fully mediated all three.

The authors’ own caveat, verbatim: “As is the case with all cross-sectional studies, we cannot determine causality, thus making the reversed hypothesis (ie, that poor sleep leads to increased binge viewing) also possible.”

The one experiment

Against all of that, there is a single randomised field experiment, and it is not about viewers at all.

Schaffner, Ulloa, Sahni, Li, Cohen, Messier, Gao and Chetty (2025), in Proceedings of the ACM on Human-Computer Interaction, recruited heavy Netflix users and randomly assigned them to keep autoplay on or turn it off. Participants supplied their own viewing histories through data subject access requests to Netflix.

OutcomeEffectp
Average daily watching-21.0 minutes0.003
Average session length-17.6 minutes0.013
Average interim time+0.407 minutes0.0009
Time between sessionsno effect reported0.697

Turning off one setting was associated with twenty-one fewer minutes of viewing a day. That is the single most actionable finding in this entire literature, and it required nothing of the viewer except a one-time toggle.

About half the treatment group said they intended to turn autoplay back on once the study ended.

Why we are not going to oversell that experiment

This is the best causal evidence anyone has. It is also small and unusual, and an article that quoted “p = 0.003” without saying so would be doing the thing this programme exists not to do.

The study analysed 76 people - 38 per arm - over an exposure period of ten to seventeen days, recruited through Prolific from US heavy Netflix users.

It reports no confidence intervals anywhere.

There is no mention of preregistration.

The analysis is bespoke. Rather than a t-test, regression, difference-in-differences or mixed model, the authors split each participant’s baseline into all possible contiguous periods of the study’s length, assembled 164 overlapping “baseline slices,” treated the spread of those slices as a distribution, converted each group’s result to a Z-score against it, and then placed the difference of the two Z-scores on a distribution of differences. Those overlapping windows are not independent draws, and reading a p-value off their spread is not a standard procedure. We could not establish that any correction for multiple comparisons was applied.

The tell is in the table. The effect with the smallest p-value, 0.0009, is an increase of 0.407 minutes - twenty-four seconds. The finding with the largest practical magnitude carries the weaker p-value. A procedure that assigns its greatest confidence to its most trivial quantity earns scrutiny.

And the finding it could not have detected

The study measured Netflix, using Netflix’s own data. It had no instrument capable of seeing where the twenty-one minutes went.

That matters, because there is good evidence the time moves. Allcott, Gentzkow and Song (2022), in the American Economic Review, randomised roughly 2,000 American smartphone users across screen-time bonus and limit treatments. Their finding on limits, verbatim: “The limit induces substitution of 12 minutes per day, so that roughly half of the FITSBY screen time that the limit eliminates moves to other apps where people had been less likely to set limits.”

Note the asymmetry they report: the bonus treatment produced no detectable substitution, with confidence intervals ruling out any substantial substitution against a 56-minute-a-day reduction. So substitution is real but depends on the design of the intervention.

There is a related caution from Lyngs and colleagues (2020) at CHI. In a six-week study of 58 university students using a browser extension, removing the Facebook news feed - the condition most analogous to removing autoplay - did not significantly reduce total time on Facebook. Only visit length fell, from 1 minute 12 seconds to 56 seconds, t(13) = 2.81, p = 0.01, d = 0.75. The condition that did reduce total time was goal reminders, not feed removal.

The numbers we will not print

Two figures dominate public writing on binge-watching, and neither can be verified.

The “73% of viewers binge” family. This traces to a Netflix press release of 13 December 2013, reporting an online survey of 3,078 US adults conducted by Harris Interactive on Netflix’s behalf via an omnibus product. No margin of error, no questionnaire, no weighting scheme, no data file, no independent report. It is also routinely miscited: in the release, 73% is the share who defined binge-watching as two to six episodes in one sitting, and separately the share reporting positive feelings about it. The figure for regularly binging is 61%.

The Netflix “Binge Scale.” Proprietary internal viewing data across more than 100 series in more than 190 countries, defining “devoured” as more than two hours a day. No methodology, no dataset, no external audit, no peer review, and the two-hour threshold is a marketing boundary rather than a construct.

Both are company-produced measurements of a behaviour that company sells. We are naming them so that they can be recognised, not using them.

What this leaves you with

If you want to watch less, change the setting rather than the resolve. The only causal evidence in this area concerns a default, and the effect was substantial. It is one small unregistered study, and it is still the best thing anyone has.

Expect some of the time to go somewhere else. The substitution research says roughly half of what a limit removes reappears elsewhere. That is not a reason to skip the setting; it is a reason not to expect twenty-one minutes back.

Treat “binge-watching” statistics as unstable. With nineteen definitions in the literature and the two most-quoted figures produced by a streaming company’s marketing department, most numbers on this subject are not comparable to each other.

Nothing in this evidence base supports treating heavy viewing as a disorder. The field’s own systematic review says the opposite, and no diagnostic manual lists it.

For the design pattern that removed the stopping cue in feeds, see infinite scroll. For what watching while doing something else costs, see second screening. For why the evening viewing matters most, see why you can’t stop scrolling in bed at night.

Limitations of this article

The entire correlational literature is cross-sectional. No longitudinal or panel study of binge-watching surfaced in our searches.

The one experiment is 76 people with a bespoke analysis and no confidence intervals. We have described its method in detail rather than quoting its p-values alone.

The field is small and its gatekeepers overlap with its authors. The systematic review carries an editorial note that it was handled by the Editor-in-Chief rather than the section editor, because an author was the section editor of the topical collection it appeared in. That was disclosed and handled properly by the journal; we disclose it too. The same Editor-in-Chief is a co-author on the meta-analysis and declares consulting relationships with pharmaceutical and gambling entities.

We did not verify one sentence in the Exelmans paper concerning the proportion of poor sleepers, because our extraction of that line was unreliable. See the withheld-figures record.

About the Author

Human Operating System

Human Operating System is a research-led publication about the human mind under digital pressure. We report what the evidence does - and does not - support.

About Human Operating System

Sources & Further Reading

7 sources

These are the sources used for this article. Where a study's limits matter to the claim, those limits are kept in the citation.

View all 7 sourcesHide sources
  1. Flayelle, M., Maurage, P., Ridell Di Lorenzo, K., Vögele, C., Gainsbury, S. M., & Billieux, J. (2020). Binge-watching: What do we know so far? A first systematic review of the evidence. Current Addiction Reports, 7(1), 44–60. Open source ↗
  2. Alimoradi, Z., Jafari, E., Potenza, M. N., Lin, C.-Y., Wu, C.-Y., & Pakpour, A. H. (2022). Binge-watching and mental health problems: A systematic review and meta-analysis. International Journal of Environmental Research and Public Health, 19(15), 9707. Open source ↗
  3. Exelmans, L., & Van den Bulck, J. (2017). Binge viewing, sleep, and the role of pre-sleep arousal. Journal of Clinical Sleep Medicine, 13(8), 1001–1008. Open source ↗
  4. Schaffner, B., Ulloa, Y., Sahni, R., Li, J., Cohen, A. K., Messier, N., Gao, L., & Chetty, M. (2025). An experimental study of Netflix use and the effects of autoplay on watching behaviors. Proceedings of the ACM on Human-Computer Interaction, 9(2), CSCW1, 1–22. Open source ↗
  5. Allcott, H., Gentzkow, M., & Song, L. (2022). Digital addiction. American Economic Review, 112(7), 2424–2463. Open source ↗
  6. Lyngs, U., Lukoff, K., Slovak, P., Seymour, W., Webb, H., Jirotka, M., Zhao, J., Van Kleek, M., & Shadbolt, N. (2020). “I just want to hack myself to not get distracted”: Evaluating design interventions for self-control on Facebook. Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, 1–15. Open source ↗
  7. Netflix. (2013, December 13). Netflix declares binge watching is the new normal [Press release]. Open source ↗
The Weekly System

The Weekly System

Join the launch list for one calm, research-led email about attention, memory, learning and the systems designed to hold your attention. No noise, no panic and no unsupported certainty.