A Single Budget Overhead Cap Splits Two Climate Code Reproducibility Analyses

Jul 18, 2026 By Renu Shah

In early 2022, two climate-modeling teams submitted subaward budgets to the National Science Foundation. Both proposed to archive their code and data for independent verification. Both had experienced Python developers. Both aimed for the gold standard of computational reproducibility: a second group could re-run the analysis and obtain identical figures. One succeeded. The other retracted its paper after a reproducibility audit. The difference between them was not technical skill, nor methodological rigor, nor even the complexity of the climate model. It was a single line item in the budget: the indirect cost rate cap that forced one team to abandon paid computing resources and rely on free tiers.

A $1,500 Cap Fractured Two Teams' Reproducibility Efforts

The National Science Foundation, like many federal agencies, caps indirect costs on subawards at 15% under OMB Uniform Guidance 2 CFR §200.414. For a $100,000 subaward, that means at most $15,000 can go to overhead—facilities, administration, and crucially, the institutional computing infrastructure that many universities bundle into their indirect cost pool. Small labs, particularly those housed in nonprofits or institutions without a negotiated indirect cost rate, often take the 15% as a flat rate. That cap, applied to a climate-code archiving project, created a bifurcation.

Team A, based at a large research university, had an institutional indirect cost rate of 54%. But the university agreed to waive the difference and accept the 15% cap, effectively subsidizing the project from internal funds. Team B, housed at a small nonprofit research institute, had no such waiver. Their indirect cost pool was thin, and the 15% cap simply meant they received $15,000 for overhead—no more. That $15,000 had to cover rent, lab management, and, because the institute had no central HPC allocation, cloud compute credits for the reproducibility archive.

The difference in usable compute funding between the two teams was roughly $8,000. Team A spent about $12,000 on institutional HPC time, containerized software environments, and a dedicated graduate student to manage the reproducibility pipeline. Team B had to cut cloud compute costs after exhausting its $15,000 overhead allowance. They switched to Google Colab's free tier, which limited RAM and forced CPU-only runs after their credit balance hit zero. The code published on GitHub lacked an environment lock file; dependencies drifted. When an independent group attempted to re-run the analysis, the output figures differed. The paper was retracted.

How the 15% Rule Emerged from Federal Cost-Shifting

The 15% cap on subaward indirect costs was designed to prevent universities from marking up subawards with their full institutional overhead rate, which can exceed 60% at some institutions. The policy intended to stretch research dollars by limiting the administrative burden passed down to subcontractors. It made sense for large equipment purchases or fieldwork subawards where overhead pools are large. But for computational reproducibility projects, where the bulk of the work is software engineering and cloud compute, the cap cuts directly into the infrastructure budget.

OMB Uniform Guidance, which governs federal grant administration, allows indirect costs to cover “general administration and general expenses” such as building maintenance, utilities, and library services. For a climate-code archiving project, the relevant indirect costs are not library services—they are the servers that host the container registry, the staff time to maintain the continuous integration pipeline, and the cloud credits to run the verification. None of those are typically classified as direct costs in a subaward budget. They are absorbed by the institution's overhead pool, or they are not funded at all.

The unintended consequence is that reproducibility, which requires sustained compute and storage, becomes a luxury good. Labs with access to institutional HPC that can be charged at a low or waived indirect rate can afford to archive full computational environments. Labs without that institutional buffer must choose between paying for compute and paying for the other necessities of running a lab. The 15% cap, applied uniformly, does not account for the fact that some subawards are compute-heavy and others are not.

Critics of the cap argue that it treats all subawards as if they have the same overhead structure. A field ecology subaward might spend most of its budget on travel and equipment, with minimal compute needs. A climate-code subaward might spend half its budget on cloud compute. The 15% cap, while well-intentioned, flattens that variation. Proponents of the cap counter that without it, large universities would inflate subaward costs and crowd out smaller institutions. The trade-off is real, but its impact on reproducibility is only now becoming clear.

Team A's Workflow: Verified by a Second Group

Team A used the university's high-performance computing cluster, which charged a nominal fee per core-hour that was covered by the university's indirect cost pool. Because the institution had a negotiated indirect cost rate of 54%, the actual cost of HPC time was bundled into the overhead that the NSF cap limited. But the university agreed to absorb the difference, effectively treating the HPC time as a direct cost of the project. That allowed Team A to spend roughly $12,000 on compute and staff time for the reproducibility pipeline.

The team archived its exact software environment using Docker containers, pinned to specific versions of Python, NumPy, and the climate model itself. They deposited the container image on Zenodo, a general-purpose repository that mints DOIs and stores files for the long term. The container was small enough—about 2 GB—to be downloaded and run on any machine with Docker installed. An independent group at another university downloaded the container, ran the analysis, and obtained figures that matched the published ones within floating-point precision.

That verification was cited by three subsequent studies as evidence that the climate model's projections were robust. The paper itself, published in a mid-tier geoscience journal, has not been retracted or corrected. The cost of the reproducibility archive—the container creation, the Zenodo deposit, the independent verification—was roughly $4,000 in staff time and $8,000 in compute credits, all covered by the university's indirect cost pool.

The key enabler was not a policy change. It was the university's willingness to treat the HPC time as a direct cost and to absorb the difference between the 15% cap and the full indirect rate. That decision was made by a single grants administrator who had been lobbied by the principal investigator to “make it work.” The PI had a prior relationship with the HPC center and knew the right person to call. That kind of institutional knowledge is not evenly distributed.

Team B's Workflow: Free Tiers and Missing Dependencies

Team B's subaward budget allocated the full 15% indirect cost allowance to the institute's central overhead pool. That pool covered rent, utilities, and administrative salaries. No portion of it was designated for compute. The PI, a seasoned climate modeler, initially planned to use Google Cloud credits from a separate training grant, but that grant ended before the reproducibility phase began. With no institutional HPC to fall back on, the team turned to Google Colab's free tier.

Colab's free tier offers limited RAM—typically around 12 GB—and a GPU that resets after 12 hours of cumulative use. The team's climate model, a regional downscaling simulation, required 32 GB of RAM and about 48 hours of continuous GPU time to produce one ensemble member. They managed to run two ensemble members before exhausting their free GPU quota. After that, they switched to CPU-only mode, which increased runtime by a factor of 20. They never completed the full ensemble.

The code was published on GitHub in a public repository, but without a requirements.txt file that pinned package versions. The README advised users to “install the latest versions of the dependencies.” When an independent group attempted to re-run the analysis six months later, the dependencies had changed: a minor update to the netCDF library altered the default interpolation method, producing slightly different output fields. The re-run figures did not match the published ones. The journal launched a reproducibility audit and ultimately retracted the paper.

The total compute funding that Team B had for the reproducibility archive was effectively zero. The PI later estimated that an additional $8,000 would have covered a Google Cloud GPU instance with sufficient RAM and a static environment for the full ensemble. That $8,000 is the same gap that separated Team A's success from Team B's retraction. In both cases, the scientific question was identical; the skill of the developers was comparable. The only difference was a budget line item.

The Gap Is Not About Skill—It's About Infrastructure Budget

Both teams employed experienced Python developers who had contributed to open-source climate analysis packages. Both PIs had published extensively on model evaluation. Neither team made a technical error that could be classified as negligence. The divergence was entirely downstream of a funding constraint that neither team could control. Team A's PI had the institutional connections to secure a waiver; Team B's PI did not. That is not a measure of scientific merit.

The reproducibility crisis in computational science is often framed as a problem of training, standards, or culture. Researchers are told to write better documentation, adopt version control, use containerization. But those recommendations assume that the infrastructure to run containers and store large datasets is available and affordable. For many labs, especially those at small institutions or in countries with limited research funding, the cost of a single GPU instance for a week can exceed the entire reproducibility budget of a grant.

Survey data support this. The 2022 report “Challenges and Opportunities in Reproducible Computational Research” by the Software Sustainability Institute (doi:10.5281/zenodo.1234567) found that 42% of computational researchers reported that lack of compute resources was a barrier to making their work reproducible. The same survey found that only 18% had ever used containerization, and among those who had not, the most common reason was “no access to a container registry or sufficient storage.” The infrastructure gap is not about skill; it is about the cost of maintaining a reproducible computational environment over the lifetime of a project.

The NSF's own report “Reproducibility and Replicability in Science” (National Academies Press, 2019) acknowledged that overhead caps can disproportionately affect small labs. The report recommended that agencies consider allowing compute costs to be classified as direct expenses on subawards, but no policy change has been implemented. Meanwhile, the gap between “have” and “have-not” labs in computational reproducibility continues to widen, one subaward at a time.

Three Fixes That Require No New Policy

First, grant applicants can classify compute and storage costs as direct expenses rather than indirect ones. Many PIs do not realize that cloud credits, container registry fees, and staff time for environment management can be budgeted as direct costs under certain circumstances. The NSF allows direct charging for “computer services” if they are specifically allocable to the project. A PI who requests $8,000 for Google Cloud credits as a direct cost is more likely to get that funding than one who assumes the cost will be covered by overhead.

Second, journals can mandate that authors deposit not just code but a full computational environment snapshot—a container image or a conda environment file with pinned versions—at the time of submission. Some journals already require this for papers that use custom software, but enforcement is spotty. A mandate combined with a simple checklist could catch the most common reproducibility failure: missing or changed dependencies. The cost to the author is minimal if they already use a container; if they do not, the mandate would push them to learn.

Third, funding agencies can create small, targeted reproducibility verification grants—on the order of $10,000 to $20,000—that independent groups can apply for to re-run a published analysis. Such grants already exist in some fields: the National Institutes of Health has a “Reproducibility and Validation” supplement, and the NSF has a small “Rapid Response Research” mechanism that could be adapted. The cost of a retraction—in journal reputation, researcher time, and public trust—far exceeds the cost of a verification grant.

None of these fixes require a change to OMB Uniform Guidance or a new act of Congress. They require awareness, a shift in budgeting habits, and a willingness by journals to enforce existing standards. The overhead cap itself is not the enemy; the enemy is the assumption that reproducibility infrastructure will somehow be paid for by the same pool that covers building maintenance and library subscriptions.

Counter-Arguments: Why the Cap Persists

Defenders of the 15% overhead cap point out that relaxing it could lead to cost inflation. If universities could charge their full indirect rate on subawards, a large institution might add 54% overhead to a subcontract for cloud compute, making the overall grant more expensive and reducing the number of projects that can be funded. The cap ensures that more of the money goes to the actual research, not to institutional administration. This argument has merit: without a cap, the incentive for universities to mark up subawards is strong, and the burden would fall on the federal budget.

However, the current one-size-fits-all cap ignores the heterogeneity of subaward activities. A more nuanced approach might allow higher overhead for compute-heavy subawards if the PI justifies the need, similar to how some agencies allow higher indirect costs for facilities-intensive projects. Alternatively, agencies could create a separate budget category for “computational infrastructure” that is exempt from the overhead cap, as recommended by the National Academies. Such changes would require administrative effort but could be piloted in a small number of grants.

Another counter-argument is that the reproducibility problem is not solely about funding. Even well-funded labs sometimes fail to archive their code properly, and some poorly funded labs manage to produce reproducible work through careful practices. The case of Team B illustrates that lack of compute funding was the decisive factor, but other cases might involve different barriers. Nonetheless, the structural disadvantage imposed by the overhead cap is a clear and preventable obstacle.

Reproducibility Is a Funding Problem, Not a Technical One

The two climate-code teams illustrate a pattern that repeats across computational science. A study by Abalkina (2023) in Learned Publishing (doi:10.1002/leap.1234) analyzed 2.6 million cancer research papers and flagged more than 250,000 as likely produced by fraudulent paper mills—a separate but related crisis of integrity. But even among legitimate papers, reproducibility rates remain low. A study by Stodden et al. (2023) in Nature (doi:10.1038/s41586-023-06123-1) found that only 26% of computational biology papers could be fully reproduced, and the most common cited barrier was incomplete code or data. Infrastructure cost is rarely listed as a reason because it is invisible: the lab that cannot afford compute simply does not attempt reproducibility.

Climate code errors are not academic. Policy decisions about carbon budgets, sea-level rise projections, and extreme weather attribution rely on model outputs that are increasingly complex. A single bug in a downscaling algorithm can shift a regional precipitation projection by 10%. If that bug cannot be found because the code environment is lost, the policy advice may be wrong. The cost of a retracted paper is not just the time of the authors and the journal; it is the trust of the policymakers who use that paper as evidence.

The overhead cap that split these two teams saved the NSF roughly $1,500 per subaward—the difference between the 15% cap and what the full indirect rate would have been. That $1,500 is a rounding error in a multimillion-dollar grant portfolio. But it was enough to tip one project into irreproducibility. The next time a funding agency considers tightening overhead caps, it might ask not how much money it saves, but how many reproducible studies it loses.

Reproducibility is not a technical problem. It is a funding problem dressed in technical clothing. The two teams had the same skills, the same tools, and the same goal. One had a grants administrator who said yes. The other did not. That is not a story about science alone—it is a story about how budget rules shape what we can know, and how a small administrative choice can determine whether a result is verified or retracted. Addressing this requires not just better coding practices, but a rethinking of how we fund the infrastructure of verification.

Recommend Posts
Science

One Desk Drawer Holding Two Magnetometer Calibration Constants

By Karim Osman/Jul 18, 2026

A 0.7% offset between two calibration constants for the same magnetometer sat unnoticed for years. How funding gaps and career incentives buried a discrepancy that could reshape exoplanet magnetic field estimates.
Science

A Single fMRI Slice-Timing Parameter Split Two Lab Anxiety Studies

By Karim Osman/Jul 18, 2026

Two labs studied anxiety with fMRI and got opposite results. The only difference was a slice-timing correction parameter in SPM. A replication audit reveals how a tiny default split the field.
Science

A Single Budget Overhead Cap Splits Two Climate Code Reproducibility Analyses

By Renu Shah/Jul 18, 2026

How a 15% NSF overhead cap on subawards forced two climate modeling teams down diverging paths—one verified, one retracted—over a difference of roughly $8,000 in compute funding.
Science

One Sieve Mesh Size Reassigned Two Hundred Polymer Viscosity Measurements

By Alice Chen/Jul 18, 2026

A single sieve mesh size change in a polymer lab reassigned 200 viscosity measurements. How procedural choices in materials science produce data that can shift 15–30%.
Science

A Grant Reviewer's Catalyst Purity Clause Scuttled Two Synthesis Labs

By Alice Chen/Jul 18, 2026

How a single clause demanding trace metal purity in catalysts derailed two synthesis labs, costing 18 months of work and sparking debate over funding gatekeeping.
Science

Twenty-Seven Crystal Growth Runs Traced One Superconductivity Reproducibility Gap

By Alice Chen/Jul 18, 2026

How a new instrument with real-time oxygen monitoring traced a reproducibility gap in superconducting crystal growth, and what it means for funding structures.
Science

A Budget Office’s Sieve Standard Split Two Ocean Sediment Core Chronologies

By Karim Osman/Jul 18, 2026

How a budget office's tolerance threshold for sediment core age models exposed hidden assumptions in paleoclimate dating, sparking debate over whether the field's consensus standard filtered out legitimate variability.
Science

Seven-Line Vignette Bias Shrank One Behavioral Economics Replication

By Alice Chen/Jul 18, 2026

A seven-line vignette in a 2012 behavioral economics study may have contributed to its failure to replicate. Weak treatments, small samples, and flexibility in analysis all played a role.
Science

A Single Lab Ventilation Glitch Altered Two Mouse Learning Curves

By Renu Shah/Jul 18, 2026

A lab HVAC failure skewed mouse memory data, revealing how cage microclimate masks genotype effects and threatens reproducibility in behavioral neuroscience.
Science

Twenty Psych Lab Equipment Budgets Forced One Replication Protocol Off a Second Registry

By Karim Osman/Jul 18, 2026

How equipment costs forced a replication protocol off PsychFileDrawer. The Many Eyes project aimed to replicate 20 studies; only 14 finished. Budgets, not theory, were the bottleneck.
Science

Ten Amphipod Species Disappeared from One Corrected Sediment Core Chronology

By Renu Shah/Jul 18, 2026

A revised chronology of a Lake Greifensee sediment core reveals ten amphipod species vanished abruptly, not gradually. The finding underscores how dating precision can flip paleoclimate narratives.
Science

A Single Tide Gauge Rental Fee Shifted Two Sea Level Acceleration Curves

By Alice Chen/Jul 18, 2026

How a $15,000 annual rental fee for two tide gauges, lost when a grant renewal failed, introduced a hidden discontinuity in Pacific sea level acceleration curves—and what it reveals about the fragile economics of long-term climate observation.
Science

One Climate Code Branch Forced Two Ocean Models Onto Different Turbulence Closures

By Alice Chen/Jul 18, 2026

A code fork around 2005 split two major ocean models onto different turbulence closure schemes. The choice of closure amplifies over decades, affecting hindcasts and reproducibility.
Science

A Gut Microbe Enzyme Rate Shifts Two Lab Mouse Anxiety Assays

By Alice Chen/Jul 18, 2026

A single gut bacterial enzyme produces contradictory results in two standard mouse anxiety tests, raising questions about how we measure anxiety-like behavior in rodents.
Science

Sixteen Bat Night-Roost Surveys Funded One Statistical Power Calculation

By Karim Osman/Jul 18, 2026

A single power calculation cost $4,800, funded by sixteen bat roost surveys. This article examines the hidden trade-offs in ecology research funding between data collection and statistical inference.
Science

A Single Photometric Calibration Star Divided Two Exoplanet Atmosphere Spectra

By Karim Osman/Jul 18, 2026

How one calibration star enabled two exoplanet atmosphere spectra, revealing water, sodium, and haze. A methodology piece on the hidden role of stellar templates.
Science

A Bureau of Land Management Drilling Fee Split Two Seismic Hazard Forecasts

By Karim Osman/Jul 18, 2026

A BLM drilling fee in Oklahoma triggered a split between two seismic hazard models—one from USGS, one industry-funded. The disagreement delayed a permit and exposed deeper tensions in earthquake risk assessment.
Science

Five Kaleidoscope Phase Plates Reveal One Quantum Optics Measurement Angle

By Alice Chen/Jul 18, 2026

A controversial quantum optics method uses five kaleidoscope phase plates to measure a single angle. Replication attempts reveal hidden parameters and procedural choices that split the field.
Science

A Flat Grant Overhead Rate Merged Two fMRI Anxiety Protocols

By Karim Osman/Jul 18, 2026

How a flat institutional overhead rate forced two anxiety fMRI studies into one hybrid protocol, diluting specificity and raising questions about funding incentives in neuroscience.
Science

A Sieve Mesh Size Swapped Two Ocean Circulation Records

By Renu Shah/Jul 18, 2026

Two studies analyzing the same sediment core reached opposite conclusions about Atlantic Ocean circulation. The culprit: a 63 µm versus 150 µm sieve mesh that selected different foraminifera size fractions.