What Visible Metrics Do to Your Decisions
Randomised experiments show visible counts change what you sample far more than what you conclude - and the effect is modest, topic-dependent and often self-correcting.

Every interface you use puts a number next to something. Likes, upvotes, view counts, star ratings, "12,400 people bought this today." The number is not describing the thing. It is describing what other people did about the thing, and it is placed there because it changes what you do next.
That much is well established by randomised experiments - real ones, with control groups, on real platforms. What is much less well established is how big the effect is, which direction it runs, and whether it survives contact with the world. The popular version of this research is considerably stronger than the research.
Note as you read: nearly every study below is on a specific platform with a specific population. There is no experiment here on a general population using a mainstream social network. That absence matters, and we flag it at the end.
The founding experiment, and what it actually was
Salganik, Dodds and Watts (2006) built a website - not a real music service, a purpose-built one they called MusicLab - stocked with 48 songs by 48 bands nobody had heard of. 14,341 participants, mostly US teenagers and young adults recruited from bolt.com, were randomly assigned either to see download counts from previous participants or to see nothing but band and song names.
The count-seeing participants were split across eight parallel "worlds," each of which accumulated its own download history independently from zero. Nothing was faked. The counts were honest; they simply had different histories.
Two findings. Popularity was much more unequal in the social-influence worlds than in the independent one (p < 0.01), and it was much more unpredictable - the same song could finish near the top in one world and near the bottom in another.
The sentence that gets quoted least and matters most is the authors' own limit on their result:
"The 'best' songs never do very badly, and the 'worst' songs never do extremely well, but almost any other result is possible."
Quality set the range. Visible history decided where in that range a song landed.
And the authors were explicit about what their study was not: "Our experiment is clearly unlike real cultural markets in a number of respects. For example, we expect that social influence in the real world - where marketing, product placement, critical acclaim, and media attention all play important roles - is far stronger than in our experiment."
What was tested: a researcher-built site, 48 unknown songs, US teenagers, 2004-05, no algorithm, no feed, no follower graph, no advertising, no social ties between participants. It is evidence that when quality is genuinely ambiguous and history is visible, outcomes are path-dependent. It is not evidence about Spotify or TikTok, and its authors did not claim it was.
What happened when they inverted the rankings
Salganik and Watts (2008) went further. They let two MusicLab worlds run to a steady state, then performed a single deception: they swapped the displayed download counts, rank-reversed - the most popular song was shown with the least popular song's count, second with second-to-last, and so on. After that one intervention, all counts updated honestly again for 9,996 further participants.
This experiment is usually retold as "quality won in the end." It did not.
What actually happened: "most songs experienced self-fulfilling prophecies, in which perceived - but initially false - popularity became real over time," and "almost all songs seem to be permanently affected by the inversion." Only the very best songs clawed back: "the success of the very 'best' songs was essentially unaffected, even though these were typically the most severely penalized."
The decisive number is the correlation between the post-inversion ordering and the original one. It started at -1 by construction, and it climbed "towards what appears to be an asymptotic limit around zero."
Not back to +1. To roughly zero - that is, uncorrelated with the ordering the market had produced on its own. The very top song returned to the top; the second-place song did not return to second place.
The authors' own verdict is the honest one: the experiment "provides some ammunition both for proponents of self-fulfilling prophecies, and also for skeptics."
The upvote experiment, and its missing half
Muchnik, Aral and Taylor (2013), in Science, ran the cleanest test of visible vote counts anyone has managed. On a social news aggregation site, 101,281 comments posted over five months were randomly assigned to be artificially up-voted (+1), artificially down-voted (-1), or left alone. Those comments were subsequently viewed over 10 million times and rated 308,515 times by ordinary users who had no idea.
The positive result is the famous one. A single artificial up-vote "significantly increased the probability of up-voting by the first viewer by 32% over the control group (P = 1.0 × 10-6)" and "increased comments' final mean ratings by 25% (P = 2.3 × 10-11)." The advantage compounded and persisted across the whole five-month window.
The negative result is almost always dropped, and it is the more interesting one. A single artificial down-vote did make further down-votes slightly more likely (0.014 versus 0.007, P = 1.1 × 10-3). But it made up-votes far more likely: 0.099 versus 0.054 (P = 1.0 × 10-30). People saw an unfairly buried comment and corrected it. The correction was strong enough that the final ratings of down-treated comments were not statistically different from the control group's.
Herding is not symmetric. On that site, people followed praise and pushed back against pile-ons.
There is a third finding that also tends to vanish. The effect was topic-dependent: significant positive herding in politics, culture and society, and business; "no detectable herding behavior" in economics, IT, fun, and general news. Whatever this is, it is not a universal law of human psychology. It is context-sensitive enough to switch off in four of seven categories on a single website.
Two corrections to how this study is usually cited. The site is not named - the paper states there are legal obstacles to disclosing it, and describes it only as similar to Digg and Reddit. It was not Reddit. And the 32% and 25% figures are different quantities: the first is the change in the first viewer's up-vote probability, the second is the change in final mean rating. They are not interchangeable.
Does the number change what you like, or only what you look at?
This is the sharpest question in the field, and the best answer comes from a reanalysis of the original MusicLab data.
Krumme, Cebrian, Pickard and Pentland (2012) separated the two decisions participants made: whether to click a song, and whether to download it once they had listened.
"Contrary to conventional wisdom, social influence is material to the first step only... The behavior of others impacts what an individual will try, but has only indirect effect on what he buys."
The download rate, conditional on having listened, was independent of the social information. The count changed what people sampled. It did not change how much they liked what they sampled.
A second reanalysis, by Lynn, Walker and Peterson (2016), found a narrow exception: popularity "can boost perceptions of a song's likeability but only for songs of lower quality" - a halo that operates on weak items and not on strong ones.
Put together, the mechanism looks much more like attention allocation than like altered judgement. That is a smaller claim than "metrics change what you think is good," and it is the one the data supports.
Informational or normative? One study separates them cleanly
There are two ways a visible count could work. It could be informational - the number is evidence that the thing is good. Or it could be normative - you want to do what others do, or be seen to.
Cai, Chen and Fang (2009) built an experiment that pulls them apart, and it is not online at all. Across 13 branches of a Szechuan restaurant chain in Beijing in October 2006, tables were randomly assigned to one of three conditions:
- Control - no plaque.
- Ranking - a plaque listing the five most popular dishes from the previous week, ranked.
- Saliency - a plaque listing five sample dishes including the top three, but not identified as popular, in random order.
Across 12,895 bills analysed, the ranking plaque raised demand for the listed dishes by 13 to 20 percent. The saliency plaque - same dishes, same visual prominence, no popularity information - produced no significant effect.
Two things make this the strongest evidence here for informational influence specifically. The saliency condition rules out mere attention. And the choice was private: nobody at another table saw what you ordered, so conformity and impression management cannot explain it. The authors also found the effect stronger among infrequent customers - exactly what you would predict if the number is working as information for people who lack it.
How large are these effects outside a lab?
Modest, and diminishing. Van de Rijt, Kang, Restivo and Patil (2014) ran randomised interventions on four real platforms simultaneously:
| Platform | Intervention | N | Control → Treatment |
|---|---|---|---|
| Kickstarter | Donated 1% or 10% of goal | 293 projects | 39% → 70% received later funding |
| Epinions | Rated reviews "very helpful" | 481 reviews | 77% → 90% got a helpful rating in 14 days |
| Wikipedia | Gave one award on user page | 521 editors | 31% → 40% received later awards |
| Change.org | Added twelve signatures | 200 petitions | 52% → 66% got more signatures |
Real effects on four live platforms. But the shape of them is the finding:
"Each additional unit increase in input yields a progressively smaller increase in output. Indeed, in each experiment an increase in input from zero to one produces a significant increase in per-unit output, whereas the additional increase in input from one to four never yields a noticeable increase in per-unit output."
Going from nothing to something mattered. Going from something to four times as much did not. The authors' own conclusion: their findings "suggest a lesser degree of vulnerability of reward systems to incidental or fabricated advantages and a more modest role for cumulative advantage in the explanation of social inequality than previously thought."
And van de Rijt (2019), in the American Journal of Sociology, went further, reanalysing prior experiments and running a new one to convergence: mismatches between popularity and quality "will usually self-correct," and explaining the durable dominance of mediocre bestsellers requires identifying additional conditions that block self-correction. Social influence on its own does not produce it.
The effect shrinks when the manipulation gets realistic
This is the most under-reported pattern in the whole literature, and it is visible inside a single paper.
Messing and Westwood (2014) tested whether endorsement counts could override partisan source preference. Study 1, with 739 US MTurk participants, contrasted stories showing 10,000+ recommendations against stories showing 0-1,000. The effect was substantial: partisan selectivity dropped from M = 0.38 to M = 0.25 (t(97) = 2.95, p = .002, d = .55), and the authors wrote that "the mere presence of endorsements reduced partisan selectivity to levels indistinguishable from chance."
Study 2 used realistic counts - 1-35 versus 150-650 - with 153 undergraduates. Participants were still significantly more likely to pick endorsed stories, but the effect was d = .08. Negligible.
Same research team, same question, same paper. A manipulation far outside the range users normally encounter produced a large effect; a manipulation inside that range produced almost nothing.
Whenever you read that visible metrics powerfully shape behaviour, check what numbers were shown. A great deal of this literature runs on contrasts nobody actually sees.
What visible metrics do to what you make
The evidence here is thinner but points somewhere specific.
Burtch, He, Hong and Lee randomly and anonymously assigned Reddit's Gold Award to 905 users' posts over two months. Recipients responded by posting more, and longer - and by producing "content that exhibited significantly greater textual similarity to their own past (awarded) content." The award pushed them toward exploitation rather than exploration. They repeated themselves.
Restivo and van de Rijt (2012) gave barnstars to 100 of Wikipedia's top-1% editors, with 100 matched controls, and found productivity up 60% (z = 3.222, p = 0.001) and subsequent awards received by 12 treated editors versus 2 controls (chi-squared = 7.681, p = 0.006). Note the population: elite volunteers, not typical users.
And Chen, Harper, Konstan and Li (2010) ran the field experiment with the most uncomfortable result. On MovieLens, showing users the median contributor's rating count produced a 530 percent increase in monthly ratings among below-median users - and a 62 percent decrease among above-median users. The metric pulled everyone toward the middle. It motivated the laggards by demotivating the leaders.
What this evidence does not support
- It does not support "metrics manipulate everyone." Muchnik found herding in three of seven topic categories. Kivetz-style universality is not what these data look like.
- It does not support treating negative and positive metrics alike. In the one large randomised test, artificial down-votes produced net correction, not net herding.
- It does not support "small early leads snowball forever." The four-platform field experiment found sharply diminishing returns, and the same author's 2019 paper argues these processes usually self-correct.
- It does not support any claim about view counts. We looked for a randomised experiment on view counts and choice and found none. We also found no published data from Meta's Instagram like-hiding trials - the company has never released the underlying results, and the only company-attributed statement we could verify is that "some people found this beneficial but some still wanted to see like counts."
- It does not support generalising from a music market to a social feed. MusicLab had no algorithm, no feed, no social graph and no advertising. Its authors said so first.
What it does support
That a visible count is not neutral information - it is an intervention placed in your path, and the best-identified thing it does is change what you sample, not what you conclude once you have sampled it. That the effect is real but modest, front-loaded, topic-dependent, and often self-correcting when a crowd can see that a judgement was unfair. And that the same number can motivate one person and deflate another, depending entirely on which side of it they land.
The practical version: treat a count as evidence about attention, not about quality. It reliably tells you what got looked at. Whether the thing was any good is a question the number was never answering.
This article is about what a visible number does to you. The opposite direction - what a feed works out about you from what you do - is covered in what your feed knows about you. The design context sits in the persuasive design guide.
A note on populations
Every experiment above tested somebody specific. US teenagers on a research website in 2004. Anonymous commenters on an unnamed news aggregator. Diners in thirteen Beijing restaurants in 2006 - the only non-Western sample here, and the only setting where choice was genuinely private. The top 1% of Wikipedia editors. 32 adolescents in an fMRI scanner viewing a simulated Instagram feed, in the one study that manipulated like counts directly. MTurk workers.
There is no randomised experiment in this set on a general population using a mainstream social platform, because platforms do not let outsiders run them and do not publish the ones they run themselves. Everything above is inference toward a setting nobody has been permitted to measure.
Human Operating System
Human Operating System is a research-led publication about the human mind under digital pressure. We report what the evidence does - and does not - support.
About Human Operating SystemSources & Further Reading
13 sourcesThese are the sources used for this article. Where a study's limits matter to the claim, those limits are kept in the citation.
View all 13 sourcesHide sources
- Salganik, M. J., Dodds, P. S., & Watts, D. J. (2006). Experimental study of inequality and unpredictability in an artificial cultural market. Science, 311(5762), 854-856. Open source ↗
- Salganik, M. J., & Watts, D. J. (2008). Leading the herd astray: An experimental study of self-fulfilling prophecies in an artificial cultural market. Social Psychology Quarterly, 71(4), 338-355. Open source ↗
- Muchnik, L., Aral, S., & Taylor, S. J. (2013). Social Influence Bias: A Randomized Experiment. Science, 341(6146), 647-651. Open source ↗
- Krumme, C., Cebrian, M., Pickard, G., & Pentland, S. (2012). Quantifying Social Influence in an Online Cultural Market. PLoS ONE, 7(5), e33785. Open source ↗
- Lynn, F. B., Walker, M. H., & Peterson, C. (2016). Is Popular More Likeable? Choice Status by Intrinsic Appeal in an Experimental Music Market. Social Psychology Quarterly, 79(2), 168-180. Open source ↗
- Cai, H., Chen, Y., & Fang, H. (2009). Observational Learning: Evidence from a Randomized Natural Field Experiment. American Economic Review, 99(3), 864-882. Open source ↗
- van de Rijt, A., Kang, S. M., Restivo, M., & Patil, A. (2014). Field experiments of success-breeds-success dynamics. PNAS, 111(19), 6934-6939. Open source ↗
- van de Rijt, A. (2019). Self-Correcting Dynamics in Social Influence Processes. American Journal of Sociology, 124(5), 1468-1495. Open source ↗
- Messing, S., & Westwood, S. J. (2014). Selective Exposure in the Age of Social Media. Communication Research, 41(8), 1042-1063. Open source ↗
- Burtch, G., He, Q., Hong, Y., & Lee, D. (2022). How Do Peer Awards Motivate Creative Content? Experimental Evidence from Reddit. Management Science, 68(5), 3488-3506. Open source ↗
- Restivo, M., & van de Rijt, A. (2012). Experimental Study of Informal Rewards in Peer Production. PLoS ONE, 7(3), e34358. Open source ↗
- Chen, Y., Harper, F. M., Konstan, J., & Li, S. X. (2010). Social Comparisons and Contributions to Online Communities: A Field Experiment on MovieLens. American Economic Review, 100(4), 1358-1398. Open source ↗
- Sherman, L. E., Payton, A. A., Hernandez, L. M., Greenfield, P. M., & Dapretto, M. (2016). The Power of the Like in Adolescence. Psychological Science, 27(7), 1027-1035. Open source ↗
The Weekly System
Join the launch list for one calm, research-led email about attention, memory, learning and the systems designed to hold your attention. No noise, no panic and no unsupported certainty.

