Seven-Line Vignette Bias Shrank One Behavioral Economics Replication
In 2012, a behavioral economics paper reported that a subtle reminder of social norms could boost cooperation. Participants who read a seven-line vignette about appropriate behavior subsequently gave more in a public goods game. The effect was modest — Cohen's d around 0.30 — but it fit a growing narrative that small primes could nudge prosocial behavior. A decade later, a multi-lab replication project found essentially no effect: d = 0.02, with only 2 of 20 labs detecting a positive result. The original claim now looks like a false positive, and the vignette itself may be partly to blame.
A Replication That Wasn't
The 2012 study, published in a respected journal, involved 148 participants divided into two groups. One group read a short paragraph describing a social norm of cooperation; the other read a neutral text. After reading, participants played a public goods game where they could contribute money to a shared pool. Those who had read the norm vignette contributed about 15% more on average, a difference that was statistically significant at p < .05.
The finding was cited dozens of times and helped inspire interventions in schools and workplaces. But around 2015, as the replication crisis gained attention, several labs tried to reproduce the result. Informal attempts failed, and a formal multi-lab replication was organized under the Many Labs framework. That project, published in 2019, found no significant effect across 20 independent samples totaling roughly 1,500 participants.
The tension between the original and the replication raises a question: why did the effect vanish? One possibility is that the original was a false positive, a statistical fluke. Another is that the replication did not capture some crucial feature of the original. But the evidence points more strongly to the first explanation, and a key suspect is the seven-line vignette itself.
The Seven-Line Vignette as a Weak Treatment
The vignette in the original study was brief: roughly 70 words describing a community where people cooperate. Participants read it once and then immediately played the game. There was no manipulation check — no question to verify that the vignette actually changed participants' thoughts or feelings. Without such a check, it is impossible to know whether the treatment had any psychological impact.
In later replications, researchers added attention checks and found that 18–25% of participants failed them. That failure rate suggests that many participants may not have read the vignette carefully or remembered it. If the treatment was too weak to register, any observed effect would be noisy and unreliable.
Weak treatments are a known problem in priming research. Many classic priming effects, such as those involving word puzzles or sentence scrambles, have failed to replicate when tested rigorously. The seven-line vignette was even briefer than those stimuli. It may have been a 'dose' too low to produce a consistent behavioral change, especially across different populations and settings.
One counterargument is that the original study used a paper questionnaire in a lab, while many replications used online platforms. Online participants are less attentive, on average. But even lab-based replications within the multi-lab project failed to find the effect, suggesting the weakness is not solely an online issue.
Sample Size and Publication Bias
The original study had 74 participants per condition. That sample size is typical for behavioral experiments of the era, but it is small for detecting a subtle effect. With 74 per group, the study had only about 50% power to detect a d of 0.30 at p < .05. That means if the true effect were 0.30, the study would fail to find it half the time. But when it did find it, the estimated effect would tend to be inflated — a phenomenon known as the winner's curse.
The replication project used a total sample of roughly 1,500, giving it over 95% power to detect a d of 0.30. The observed d of 0.02 is consistent with a true effect of zero. A Bayesian analysis of the combined data strongly favored the null hypothesis, with a Bayes factor of about 12.
Publication bias likely played a role. In 2012, pre-registration was rare. Researchers could analyze data in multiple ways — excluding outliers, controlling for covariates, choosing different dependent variables — and report only the analysis that yielded significance. The original study did not pre-register its analysis plan, so the degrees of freedom were high. The replication, by contrast, was a registered report with a fixed plan, reducing flexibility.
Some defenders of the original argue that the replication may have differed in subtle ways, such as the exact wording of the vignette or the instructions for the game. But the replication team coordinated with the original authors to match the materials as closely as possible. Small variations are inevitable across labs, but if the effect were robust, it should have appeared in at least a few of the 20 samples.
Context Sensitivity of Behavioral Primes
Behavioral primes are notoriously context-dependent. A subtle cue that works in one culture, language, or setting may fail in another. The original study was conducted in a Western university lab with paper questionnaires. The replications spanned 20 labs across multiple countries, some using online platforms. The shift from paper to screen may have altered how participants engaged with the vignette.
Even within the same lab, timing matters. The original data were collected in 2010–2011, before social norms around cooperation had shifted. The replications took place between 2015 and 2018, after several high-profile failures to replicate priming effects had been published. Participants may have become more skeptical or less susceptible to such primes.
One meta-analysis of similar norm-priming studies found that effect sizes declined over time, a pattern consistent with both genuine context sensitivity and publication bias. The decline was steepest for studies with small samples and weak treatments, exactly the profile of the 2012 vignette.
Some researchers argue that context sensitivity is a feature, not a bug, of behavioral science. Primes work only when they are novel and the participant is unaware. Once a prime becomes widely known, its effect may vanish. But that explanation is hard to test, and it risks making the claim unfalsifiable.
What the Replication Actually Found
The multi-lab replication, published in 2019 as a registered report, found an overall effect of d = 0.02 with a 95% confidence interval ranging from −0.05 to 0.09. That interval includes zero and excludes the original estimate of 0.30. Only 2 of the 20 individual labs found a statistically significant positive effect, and those had small samples. The remaining 18 found null or negative results.
A Bayesian meta-analysis of the replication data yielded a Bayes factor of 12.3 in favor of the null hypothesis, meaning the data were about 12 times more likely under a model with no effect than under a model with an effect of the size originally reported. Sensitivity analyses that varied the prior did not change this conclusion.
The replication also tested whether the effect might be moderated by demographic variables such as gender, age, or nationality. None of these moderators were significant. The null result was consistent across the board.
Some critics noted that the replication used a slightly different public goods game — a linear one rather than a step-level one — but the original authors had agreed that the linear version was a valid test. Moreover, a separate direct replication using the exact original game also failed to find an effect.
Lessons for Designing Replicable Studies
The case of the seven-line vignette offers several lessons for researchers. First, ensure that the treatment has a measurable impact. A manipulation check — a simple question about whether the vignette changed participants' thoughts — would have revealed whether the prime worked as intended. Without it, the original study could not distinguish between a real effect and noise.
Second, conduct a power analysis before data collection. The original study's sample size was adequate only for detecting medium-to-large effects. For subtle primes, larger samples are needed. The replication's sample was 20 times larger, giving it the precision to detect even small effects if they existed.
Third, pre-register the analysis plan. Pre-registration reduces the risk of p-hacking and selective reporting. The original study had no pre-registration, while the replication did. That difference alone may account for some of the discrepancy.
Fourth, report effect sizes with confidence intervals, not just p-values. The original reported a p-value just below .05 but did not provide a confidence interval for the effect size. A confidence interval would have shown the imprecision of the estimate.
Finally, consider the strength of the treatment. A seven-line vignette is a weak intervention. Researchers should pilot-test their stimuli to ensure they produce the intended psychological change. Weak treatments amplify noise and reduce replicability.
The Slow Path to a Settled Verdict
Seven years elapsed between the original publication and the meta-analytic replication. That is a long time for a single finding to remain unresolved. During those years, the field of behavioral economics underwent a methodological reckoning. The replication crisis spurred reforms such as pre-registration, registered reports, and larger sample sizes. The seven-line vignette study was an early target of that reckoning.
Today, the consensus among meta-analysts is that the original claim is likely a false positive. The evidence from 20 labs, combined with Bayesian analysis, is strong. But some researchers remain cautious, noting that absence of evidence is not evidence of absence. The true effect, if it exists, is certainly smaller than originally reported — possibly too small to be practically meaningful.
The case is a small but instructive example of how cumulative evidence works. No single study settles a question. The original was not fraudulent; it was simply underpowered and overinterpreted. The replication corrected the record, but it took years and the coordinated effort of many labs. The field is now better equipped to avoid such false positives, but the slow pace of self-correction remains a challenge.
As one researcher put it, 'Science is a process, not a product.' The seven-line vignette story is a reminder that even a well-intentioned finding can be misleading, and that the path to a settled verdict is rarely straight.
Broader Implications for Behavioral Science
The failure of the seven-line vignette to replicate is not an isolated event. It echoes similar patterns in other domains of behavioral science, such as social priming and ego depletion. For instance, the classic 'elderly walking slower' priming effect, in which participants exposed to words related to old age subsequently walked more slowly, has also failed to replicate in large-scale efforts. Similarly, the idea that self-control is a depletable resource — the ego depletion effect — has been called into question by a multi-lab replication that found a near-zero effect. These cases share common features: small original samples, weak manipulations, and a publication environment that favored surprising results.
One lesson from these failures is that the field's reliance on small, underpowered studies has produced a literature with many false positives. The seven-line vignette is a textbook example. Its effect size of d = 0.30 was plausible but imprecise, and the original study lacked the statistical power to reliably detect it. The replication, with its larger sample and pre-registered protocol, provided a more trustworthy estimate. The discrepancy between the two is a reminder that single studies, especially small ones, should be interpreted with caution.
Another lesson is the importance of direct replication. Many researchers argue that conceptual replications — studies that test the same idea using different methods — are more valuable than direct replications. But conceptual replications introduce new variables that can obscure failures to replicate. The seven-line vignette case shows that direct replication, when done carefully, can provide clear evidence against a false positive. The multi-lab approach, with its coordinated materials and analysis plan, is a model for how to conduct such replications.
Some critics of the replication movement argue that too much emphasis on replication stifles innovation. They worry that researchers will avoid risky, creative studies for fear of failing to replicate. But the seven-line vignette case suggests the opposite: by identifying false positives, replication efforts can clear the way for more robust findings. The original study's claim was not just wrong; it was distracting. Researchers who built on it wasted time and resources. A more careful approach, with pre-registration and larger samples, would have produced a more reliable result from the start.
Practical Steps for Researchers
For researchers designing similar studies, the lessons are clear. First, pilot-test your manipulation. A simple check — asking participants what they thought about after reading the vignette — can reveal whether the prime had the intended effect. Without such a check, you cannot be sure that your treatment is working.
Second, plan your sample size based on a power analysis. For small-to-medium effects, you need hundreds of participants, not dozens. Online platforms like Mechanical Turk or Prolific make large samples feasible and affordable. There is no excuse for underpowered studies in the modern era.
Third, pre-register your analysis plan. This does not mean you cannot explore your data; it just means you separate confirmatory analyses from exploratory ones. Pre-registration increases transparency and reduces the risk of p-hacking.
Fourth, report effect sizes with confidence intervals. A p-value tells you whether an effect is statistically significant, but it does not tell you how large or precise the effect is. Confidence intervals provide that information and help readers assess the reliability of the finding.
Finally, consider the strength of your treatment. If your manipulation is a brief text or a subtle cue, ask yourself whether it is strong enough to change behavior. If not, you may need a more powerful intervention, or you may need to accept that your study is better suited for detecting large effects.
Conclusion
The seven-line vignette study is a cautionary tale about the dangers of weak treatments, small samples, and publication bias. Its failure to replicate does not mean the original authors were dishonest; it means the original evidence was weaker than it appeared. The replication effort, involving 20 labs and a pre-registered protocol, provided a more reliable estimate and corrected the scientific record.
The case also illustrates the value of methodological reforms. Pre-registration, registered reports, and large-scale collaborations are not bureaucratic hurdles; they are tools for improving the credibility of science. The field of behavioral economics has embraced many of these reforms, and the result is a more robust literature. But the process is ongoing, and new challenges will arise.
As the researcher said, science is a process, not a product. The seven-line vignette story is a reminder that even a well-intentioned finding can be misleading, and that the path to a settled verdict is rarely straight. But with careful methodology and a commitment to self-correction, the path becomes clearer.