The PDN library team has reviewed the evidence on seven prosocial ranking approaches.

When you open an app like Instagram, TikTok or Facebook, there are thousands of potential posts that could appear in your feed. Feed-ranking algorithms decide which of those thousands of eligible posts you actually see, and for most major platforms that decision has been optimized primarily for engagement.
A growing body of research (see Beknazar-Yuzbashev et al., 2022 and Cunningham et al., 2025 for an overview) finds that engagement-based ranking systematically favors content that provokes outrage, moral condemnation, and hostility toward political opponents — which raises the obvious question of whether ranking can be tuned to produce different outcomes.
Until recently that question was hard to test, since ranking happens inside the platform’s systems. A number of new studies have changed that, including two that use browser extensions to re-rank the live feeds of consenting participants. The PDN library team has reviewed the evidence on seven distinct prosocial ranking approaches based on that research, which we summarize below. (More info on how our ratings work.)
A note on how we report results: effect sizes follow the metrics used in each paper, and all effects are statistically significant unless we state otherwise.
Rating: Validated
What it is: A feed-ranker that demotes content expressing hostility (e.g., insults, identity attacks, scapegoating, fearmongering, moral outrage, and antidemocratic attitudes), in particular toward other groups. Some versions pair the downranking of toxic content with the upranking of content showing markers of civility.
Evidence it works: Two field experiments, both conducted on live feeds around the 2024 U.S. election, found that downranking toxic content reduces hostility toward political opponents. Piccardi et al. (2025) ran a ten-day browser-extension experiment with X (formerly Twitter) users, downranking posts high in antidemocratic attitudes and partisan animosity ("AAPA") from roughly 10% of a participant's feed to about 1%. After a week, those participants felt about 2 degrees warmer toward counter-partisans on a 100-point feeling thermometer compared to a group with no changes to their feed. While not noted in their paper, authors report to us that effect size is equivalent to 0.09 standard deviations. Stray et al. (2026), a larger six-month experiment across Facebook, Reddit, and X, tested an algorithm that used Google Jigsaw classifiers to demote antisocial content while upranking civil content; it reduced partisan animosity by roughly 0.04 standard deviations. Notably, a separate arm of that same study that only upranked civil content produced no significant effect, which points to the downranking component as the active ingredient.
Limitations: Both studies ran during an unusually turbulent stretch of a presidential campaign, and Piccardi et al. recruited only X users whose feeds already carried at least 5% political content, so effects may be concentrated among politically engaged users in tense moments; or it may be that the politically engaged tend to be more polarized and elections can increase emotions, so that may be the desirable timing and audience. Downranking also reduced total time on platform in the X study, though those participants returned as often as before, retweeted as much, and engaged more with the posts they did see.
Rating: Likely
What it is: A feed-ranker that inserts posts from a set of credible news outlets spanning the ideological spectrum, increasing users' exposure to professional reporting on public affairs.
Evidence it works: Stray et al. (2026) tested this as one arm of their six-month field experiment across Facebook, Reddit, and X during the 2024 electoral cycle. The algorithm inserted news matched to users' prior interests, drawn from a list of credible outlets weighted across the American political spectrum. It significantly reduced partisan animosity by around 0.04 standard deviations, the largest reduction of any algorithm tested in that study.
Limitations: The evidence rests on a single field experiment. The intervention also requires a platform to decide which outlets count as credible and how to weight them across the spectrum, an editorial judgment that is difficult to make at scale. Note: this differs from Increase Exposure to Counter-Partisan Content (below) in that it adds balanced professional coverage rather than specifically exposing users to an opposite side’s perspectives or sources.
Rating: Tentative
What it is: A feed-ranker that promotes content valued by people on opposite sides of a political divide, rather than content that generates the most engagement overall.
Evidence it works: Brady et al. (2025) tested a bridging algorithm in a simulated feed environment (YourFeed), ranking posts higher when they drew engagement from both Democrats and Republicans and lower when they appealed to only one side. Compared with participants who scrolled engagement-ranked feeds, those in the bridging condition reported lower intentions to post content blaming the outgroup or praising the ingroup (d = 0.20–0.22 versus a passive-engagement feed, d = 0.27–0.29 versus an active-engagement feed). Additional analysis suggests the mechanism runs through norm perception, or people’s sense of what is typical behavior; engagement-based feeds inflated beliefs about how common and acceptable polarizing content was, which in turn increased willingness to post it. In the field, Stray et al. (2026) tested a related algorithm ("diverse approval") that used GPT-4o to predict which posts both Democrats and Republicans would find valuable. That version produced no significant reduction in affective polarization and significantly lowered engagement with the platforms.
Limitations: The significant evidence comes from a simulated environment, not a real platform, and it measured what people said they would post,rather than actual behavior. The two studies also defined cross-partisan approval quite differently. One used engagement data from Democrats and Republicans, the other from LLM models to predict what both sides would like. Because of that difference, the lack of a positive result in the field test is hard to interpret. It could mean that bridging-based ranking doesn’t work in the real world or just that the AI-prediction version doesn’t. Future work should adjudicate between these approaches and test them against current platform defaults.
Rating: Mixed
What it is: Increasing how much content users see from the other side of the political spectrum, either by adjusting ranking or by nudging users to follow cross-party accounts.
Evidence it works: Four field studies point in different directions. Levy (2021) offered Facebook users the chance to "like" out-partisan news accounts. Those given that option later expressed reduced dislike of counter-partisans than users who were not. Bail et al. (2019) paid Twitter users to follow a bot retweeting counter-partisan content and found a backfire effect. Rather than softening, participants became more entrenched in their own views. Nyhan et al. (2023), run in collaboration with Facebook, approached the question from a different direction, reducing like-minded sources in users' feeds by 17 percentage points, and found no change in attitudes toward out-partisans or in ideological extremity. Casas et al. (2023), designed in part to reproduce the backfire effect in Bail et al., incentivized participants to read ideologically extreme out-party news for twelve days and found no polarizing effect.
Limitations: The divergent findings can in part be explained by the different outcomes measured; Bail et al. looked at changes in issue polarization, while Levy focused on “affective” polarization (i.e. feelings about counter-partisans) and the two other studies measured both dimensions. One other relevant distinction between the Levy and Bail studies is the quality of what users encountered: Levy, the study that reduced animosity, used four partisan but reputable outlets, while the Bail et al. bot (the one with the backfire effect) drew from over 4,000 political accounts of varying quality. Given the risk identified in Bail et al. (2019), we advise limiting the inclusion of out-partisan content to more reputable, non-extreme sources when increasing cross-partisan exposure.
Rating: Emergent
What it is: A feed-ranker that upranks posts and comments showing markers of civility, such as reasoning, compassion, respect, curiosity.
Evidence it works: Stray et al. (2026) used Google Jigsaw's Perspective API to score posts for civility and upranked the highest-scoring content as one arm of their six-month field experiment. The result was a non-significant reduction in partisan animosity, with likewise no significant effect on user experience, social trust, intergroup empathy, or political knowledge.
Limitations: The evidence amounts to a single null result, one study that showed no statistically significant effect. Read alongside Downranking Toxic Political Content, where a combined downrank-and-uprank algorithm did work, it suggests that removing hostile content may be doing the meaningful work, and that elevating civil content on its own may not be enough. That said, this is a relatively low-cost change for platforms, and one null finding is not evidence of failure.
Rating: Emergent
What it is: A feed-ranker that promotes posts in which members of one political group express views usually associated with the other (e.g., a conservative account voicing a liberal position), on the logic that seeing the other side as less uniform reduces animosity toward it.
Evidence it works: Stray et al. (2026) used GPT-4o to identify and promote stereotype-challenging posts in one arm of their field experiment. The reduction in partisan animosity was not statistically significant. Two secondary effects did reach significance, however: users reported a worse experience on the platform (-0.06 standard deviations), while the condition also produced the largest increase in engagement of any algorithm tested in the study.
Limitations: This the first large-scale field test of a long-standing idea, and the absence of a significant effect on partisan animosity leaves it unresolved rather than disproven. The pairing of higher engagement with worse reported experience is worth pondering for platforms weighing this approach: the content appears to hold attention, though users don’t enjoy it.
Rating: Unlikely
What it is: Presenting posts in the order they were published, with no engagement-based ranking at all.
Evidence it works: Guess et al. (2023) ran a three-month field experiment with more than 20,000 Facebook and Instagram users, randomly assigning some to a reverse-chronological feed and others to the platforms' default. Removing the engagement algorithm reduced engagement, as expected, but had no discernible effect on polarization or self-reported political participation.
Limitations: The comparison was against Meta's default ranking as it stood in late 2020, and ranking systems differ across platforms and over time. This therefore should be seen as a test of chronological feeds versus Meta’s 2020 engagement algorithm, rather than a general verdict on chronological feeds. It should also be noted that platform field experiments rarely detect changes in survey-measured attitudes even when behavior shifts, so the lack of effect may partly reflect what the study chose to measure. Still, the finding fits a broader pattern here: simply removing engagement ranking does not appear to deliver prosocial outcomes on its own, whereas ranking changes that actively demote hostile content do.
Much of this evidence comes from one study. Stray et al. (2026) contributes to five of the seven entries above and is the sole source for three of them. That study is unusually ambitious — six months, three platforms, and multiple algorithms tested head-to-head against a common control, which is exactly the kind of test the field has been missing. But it means our ratings are less independent of one another than the entries might suggest, and a replication that came out differently could change several of our design intervention ratings at once.
"Bridging" means at least two different things. The term is used loosely across this literature. Sometimes it describes algorithms that promote civility, harmony, and constructive tone in general; sometimes it refers specifically to algorithms that identify content with demonstrated appeal across group lines. We treat these separately — the first sits closer to Upranking Civility, the second under Bridging-Based Ranking — because the evidence diverges and the way platforms would implement them differs. A platform pursuing the first needs a way to classify the tone of the content; a platform pursuing the second needs a way to measure or predict cross-group approval, a harder problem that current studies solve in distinct ways. When you encounter the term, it is worth asking which version of “bridging” is meant.
Where the research is heading. Three developments seem most consequential. (1) The browser-extension method has effectively opened ranking research to people outside the platforms, and we expect the volume of independent field experiments to keep growing. (2) Large language models are lowering the cost of classifying content along dimensions that were previously impractical to measure at scale, widening the range of signals a platform could act on. (3) What is most needed is research that measures a wider range of outcomes in more settings. Nearly all of the research here tests partisan animosity among American users during an election year. Whether these ranking changes affect wellbeing, misinformation exposure, or everyday conversation quality, and whether they behave the same way in other political environments or outside a U.S. campaign season, remain largely open questions.
We should also note that this assessment is based on publicly available research. We assume platforms have tested ranking changes internally at a scale no independent team can match (see for context KGI’s report on Better Feeds), but those results remain inaccessible. (Naturally, we invite platforms to share that research with us and the wider public.) As new research becomes public, we'll do our best to update our assessments above. If you know of studies — or interventions — we've missed, let us know!.
Hugo (he/him) is a PhD candidate in Political Science at the University of Western Ontario. His research is focused on the formation and diffusion of political beliefs, with a special interest in digital spaces and the connections between lifestyle and politics.
The Prosocial Design Network researches and promotes prosocial design: evidence-based design practices that bring out the best in human nature online. Learn more at prosocialdesign.org.
A donation for as little as $1 helps keep our research free to the public.