The NIH and the White House have both signaled interest in a policy requiring biomedical researchers to report null results, not just exciting positive results.
This is a great idea that could solve multiple problems in one fell swoop:
1) It would help reduce publication bias by informing the world about what doesn’t work;
2) It would incentivize researchers both to take on harder research problems (knowing that a null result is acceptable) and to do less p-hacking in order to achieve a positive result.
Below, we offer several examples of how such a requirement could play out. At the same time, a national-level policy would have to be carefully written, or else it would create massive confusion. After all, there are many studies in biomedicine that don’t really fit the model of “positive results vs. null results” in the first place.
Just to take one example: A recent research article in Nature examines the structure of cytoplasmic lattices in mammalian eggs. These lattices were first discovered in the 1960s and are important to embryonic development. Here’s a diagram from a Cell paper from 2023 (that paper said that these structures are still “poorly understood” and “enigmatic” despite being first discovered 60 years ago!).
So, the recent Nature paper uses cryo-electron microscopy and AI modeling to look more deeply into these “enigmatic” structures in mouse eggs. The researchers found out that these structures “are formed by at least 13 different proteins” that assemble into a megadalton-scale complex. They then go into much more detail about the exact proteins involved, what they seem to be doing, and how they are structured, and point out that this kind of work could help us get a better understanding of what leads to “infertility and developmental defects.”
This paper, in other words, is like many papers in biomedicine: Scientists take a powerful new tool or technique (here, cryo-electron microscopy) and point it at something that isn’t fully understood yet, and then basically write up, “here’s all the cool stuff we saw.”
It’s not clear what a “null result” would even be for that type of study. It’s not as if a finding that there were 5 proteins would be a “null result” compared to 13 proteins. Nothing here is a “yes or no” question, where the answer could be “no,” as there is no prespecified hypothesis.
Moreover, many biomedical labs that have “failures” or “null results” on a daily basis. Maybe a lab technician made a mistake in pipetting, but it will be corrected tomorrow – is that a “null result” that needs to be written up? We don’t want to inadvertently create the largest bureaucratic burden ever.
A “report all results” policy
Instead, let’s think about a “report all meaningful results” policy. Not “report all the times where an accident or misunderstanding occurred” or where “we just needed to let the solution sit until Thursday instead of Wednesday.”
A “report all results” policy makes the most sense (and would be most easily implemented) where studies have a clear “yes or no” conclusion as to some particular hypothesis.
Clinical trials are the most obvious example, and at least on paper, NIH-funded clinical trials are already subject to a report-all-results rule. Since 2017, NIH policy has required all NIH-funded clinical trials to be registered on ClinicalTrials.gov within 21 days of enrolling their first participant and then, within a year of completion, to submit their results there.
In other words, the federal government has already accepted the core principle that when public money funds a study designed to test whether an intervention works, the public record should reflect what happened, and specifically whether the result is positive, null, mixed, or unfavorable. That policy is not always enforced as rigorously as it should be. A 2022 audit by HHS’s Office of Inspector General found that 37 of the 72 NIH-funded trials it reviewed had not complied with federal results-reporting requirements. But it is the right approach.
And there were good reasons for such a policy. The best-known case remains a 2008 New England Journal of Medicine study by Erick Turner and his colleagues about antidepressant trials submitted to the Food and Drug Administration. Turner and his coauthors compared the full set of FDA-registered studies with the published literature. They found a stark asymmetry: out of 38 positive trials, all but 1 were published in the medical literature. Out of 36 negative or questionable studies, only 3 were published in a way that matched the FDA’s verdict. Another 11 were published in a way that presented the results as positive, while the rest went unpublished!
In short, in a case where we had the full denominator of completed studies, we could see that the academic literature was systematically distorted in a way that made antidepressants appear far more effective than the evidence actually supported. Any human being, or AI tool, who reviews the medical literature on those antidepressants would be wildly misled.
Think how much worse the situation is in all the other areas of scientific literature where there is no FDA that requires high standards for study design and then independently checks up on the data and statistics afterwards. There is no FDA for mouse research, or for economics and psychology, etc.
In one case where we do have the full denominator (an NSF-funded experimental platform), we can see that in survey experiments, 96% of studies with strong results were written up while “nearly two-thirds of the social science experiments that produced null results . . . were simply filed away.”
Thus, it is time to expand the “disclose all results” requirement to other areas that resemble clinical trials, i.e., the study is designed to answer a predefined question that has a “yes or no” answer. The case is stronger still if the study might influence downstream decisions by researchers, funders, regulators, firms, or clinicians. Where those conditions are present, a report-all-results requirement is a no-brainer.
[In principle, we would support support a “disclose all results” requirement for basically all areas of research, but it would be unenforceable in many fields where there just isn’t any way to track what a given lab is doing on a given day and whether it came up with “results” that were subject to the requirement in the first place.]
The most obvious area to expand here would be preclinical efficacy research. For example, a researcher might test a molecule, compound, biologic, or device in an animal model to see whether it improves a particular outcome, such as tumor size, infarct volume, survival, motor performance, behavior, biomarker levels, or some other endpoint that is meant to show that the treatment might have therapeutic value.
In studies like that, the logic is very close to that of clinical trials. If the point is to test whether an intervention works, and whether it’s worth carrying forward to human trials, then null results are an essential part of the scholarly record.
Stuart’s friend Malcolm Macleod and his collaborators have spent years showing why this matters. In a widely-cited analysis of animal stroke studies, he and his colleagues found that the published literature on stroke treatment wildly overstated the possibility of effective treatments.
To be sure, NIH itself has responded to such concerns by emphasizing rigor and reproducibility and by endorsing stronger principles for reporting preclinical research (e.g., the ARRIVE guidelines). But that sort of guidance is not the same thing as a mandatory reporting rule. If anything, the Macleod-type literature shows that preclinical efficacy studies are exactly the sort of domain in which a report-all-results requirement would make sense.
The ALS field is a sad and vivid example. For years, a particular mouse model generated a long list of purported therapeutic “successes.” Yet almost none of those successes led to effective human treatments. Sean Scott and colleagues at the ALS Therapy Development Institute analyzed more than 5,000 mice and showed that the standard model was noisy, underpowered in many ordinary uses, and likely to produce exaggerated results. Michael Benatar likewise documented the disconnect between interventions that appeared beneficial in mice and their failure in human ALS trials.
One of us (Stuart Buck) has met Jamie Heywood, who founded the ALS Therapy Development Institute in a desperate attempt to cure his brother’s ALS. He was still deeply upset over “wasting all of the $20 million dollars I raised.” Literally none of the studies replicated, and his brother died.
We all know that animal models are often imperfect. The point is that an entire preclinical literature can look much more encouraging than it really is when weakly designed studies, flexible analyses, and selective publication all point in the same direction. A report-all-results requirement for preclinical efficacy studies would not solve every problem, but it would at least make the evidentiary record more complete and less biased.
A second area where the same “disclose all results” rule should apply is research validating biomarkers and surrogate endpoints. NIH funds a large amount of work in this area. A search on NIH RePORTER for “validate biomarker” returns over 10,000 results, and a quick glance at the results shows that they seem relevant.
Here’s the problem: If mostly (or only) the positive validation studies are published, while null or disappointing findings are filed away never to be seen again, the field will be misled in much the same way that drug efficacy fields are misled by selective reporting. We should expect any biomarker and surrogate endpoint study to report all results, especially when it involves prespecified endpoints and is meant to lead to meaningful decisions that might affect patient care.
A third area is genetic association research. The literature in psychiatric genetics spent decades plagued by spurious candidate-gene associations and publication bias that vanished upon replication or re-examination. As with antidepressants, a large part of the problem was selective reporting: positive associations were published, null ones were not, and the cumulative literature was likely wrong. If we had a report-all-results requirement for NIH-funded genetic association studies—particularly those with prespecified hypotheses and defined primary endpoints—we could help fix this problem.
A fourth area is safety and toxicology research. The question is a bit more complicated here, because it is often the case that exciting negative findings can be more likely to get published than null findings. Nonetheless, the ultimate question is the same: we deserve to see all the results, not just the ones that are cherry-picked for being “interesting.”
A fifth field is protein-ligand binding research, where the overwhelming majority of tested compound–target pairs produce null results. That is, most compounds do not bind meaningfully to a given target. Yet the published literature is dominated by positive binding events. Any human or AI who reads that literature will come away with a distorted view in which every candidate looks more promising than it actually is.
These areas all share the features that make mandatory reporting both essential and possible: they involve an empirical question with an answer that is basically “yes or no,” and they lead to results that will shape decisions by researchers, funders, regulators, or clinicians. Our current rules require disclosure of all results for clinical trials, and we should expand those rules to new areas that have the same features.
For studies of the kind described above, NIH should make funding contingent on preregistration and results reporting. Researchers should be required to publicly register a study plan for each covered study, linked to its grant, before data collection—or, for studies using existing data, before the planned analysis. The plan should detail their research question, the outcomes they intend to measure, and how they will analyze the results. Researchers should be free to update their plans as long as they report when and why, preserve the original registration, and clearly identify changes made after examining the results.
Upon completion of a study, researchers should be required to share results regardless of whether they are published in a journal. This could take the form of a public report in an NIH-recognized repository detailing results for all prespecified outcomes, including null and inconclusive findings. Studies that stopped early would be required to report why and share any available findings. NIH could set a reporting deadline of one year after study completion, matching the deadline for clinical trials.
Research institutions could be required to list covered studies and link to their registrations and results in their regular grant reporting, enabling NIH to audit them for completeness and consider poor adherence in future funding decisions.
A final point: Across drug discovery, target identification, biomarker development, and translational research, AI tools are increasingly being used to review and synthesize the published literature, so as to guide decisions about what to pursue next. There is immense hype about the potential for AI here, but AI tools are only as good as the data on which they are trained. The case for mandatory null-result reporting was already strong, but it is getting even stronger now. If AI tools don’t have access to all of the null results that have been hidden, they will be systematically misleading. Garbage in, garbage out.
A carefully tailored policy to “disclose all results” could have a huge impact on the future of biomedical research, and whether that research translates into benefits for real-world patients rather than just benefiting the publication records of individual researchers.




