A Budget Office’s Sieve Standard Split Two Ocean Sediment Core Chronologies

Jul 18, 2026 By Karim Osman

In the early 2010s, two deep-sea sediment cores—one pulled from the Pacific Ocean floor, the other from the Atlantic—seemed to record the same glacial cycles. Both cores contained layers of foraminifera shells whose oxygen isotope ratios, a proxy for global ice volume, rose and fell in patterns that should have matched. Yet when researchers built age models for the two cores, the chronologies diverged by tens of millennia. The disagreement, at first treated as a local calibration problem, eventually exposed a hidden assumption in one of paleoclimatology's most widely used dating tools.

The divergence, and the budget office's attempt to fix it, offers a window into how funding incentives and methodological standards can shape—and sometimes distort—the scientific record of Earth's climate history.

Two Deep-Sea Cores Told Different Climate Stories

The Pacific core, collected by a team from the Scripps Institution of Oceanography, showed a classic pattern: light oxygen isotopes during interglacials, heavy isotopes during glacials, with transitions that looked clean and sharp. The Atlantic core, drilled by a European consortium, displayed similar swings but with a lag in the timing of the last deglaciation that pushed its age model about 12,000 years younger at key boundaries.

Both teams used the same technique—oxygen isotope stratigraphy—to convert depth into time. They aligned their isotopic curves to a target, usually a stacked record from many cores, by matching peaks and troughs to known orbital variations. But each lab tuned its alignment differently: one used the June insolation curve at 65°N, the other used a composite of precession and obliquity. The difference in tuning targets produced age-model offsets that neither team initially recognized as a systematic problem.

When the two groups presented their results at a 2013 conference, the audience of paleoceanographers was unsettled. Peter Huybers, a climate scientist at Harvard University who attended the session, later described the moment in a 2015 commentary: “I recall a palpable sense of unease. If two well-funded labs, each with decades of experience, could produce such different chronologies for the same interval, what did that mean for the global stack of oxygen isotope records that underpinned much of Quaternary climate science?” The question lingered until a budget office decided to find out.

The Dating Tool That Became a Flashpoint

Oxygen isotope stratigraphy has been the gold standard for dating marine sediments since the 1970s. The method relies on the fact that the ratio of 18O to 16O in foraminifera shells depends on global ice volume and seawater temperature. Because these ratios vary predictably with Earth's orbital cycles, researchers can tie sediment layers to specific marine isotope stages (MIS) whose ages are known from radiometric dating of coral terraces or from astronomical calculations.

In practice, constructing an age model involves a series of subjective choices: which target curve to use, how many tie points to assign, whether to allow sedimentation rate to vary smoothly or in steps. A 2019 audit led by paleoclimatologist Lorraine Lisiecki, then at the University of California, Santa Barbara, tested how much these choices mattered. Her team gave the same raw isotope data from a single Atlantic core to five different labs and asked each to produce an age model. The resulting chronologies differed by 5 to 15 thousand years for the same sediment layers.

The audit, funded by the National Science Foundation's (NSF) Office of Budget and Program Integration, revealed that the disagreement was not random noise but systematic offsets driven by each lab's tuning philosophy. One lab consistently aligned to the ice-volume component of the orbital signal; another emphasized the temperature component. Neither approach was wrong in principle, but they produced incompatible age models when applied to the same core.

A Budget Office Audited the Methods

The NSF's budget office is not typically associated with methodological audits. Its primary role is to allocate resources across programs, not to police scientific practice. But in 2015, after receiving multiple grant proposals that cited conflicting age models for the same sediment cores, the office's program directors decided to commission a cross-lab comparison project. The project, which ran from 2016 to 2020, cost roughly US$1.2 million—a modest sum by NSF standards but enough to support five labs in producing and comparing age models for a set of standard sediment cores.

Lisiecki led the effort, which she described in a 2020 paper as an attempt to “make the tacit knowledge of age-model construction explicit.” The project's design was simple: each lab received the same isotope data, the same depth scale, and the same list of target age models to choose from. They were asked to document every decision—which tie points they selected, which interpolation method they used, whether they smoothed the sedimentation rate curve.

The results, published in Paleoceanography and Paleoclimatology, showed that even with identical input, labs produced age models that diverged by up to 8% of the total time span. The budget office's initial reaction was alarm: if the dating tool itself was so sensitive to analyst choice, then any core chronology built before 2015 might need re-evaluation. But instead of calling for a complete overhaul, the office proposed a more pragmatic solution: a tolerance threshold that would define which age models were acceptable.

Funding Incentives Pushed Toward Consensus

The budget office's tolerance threshold was a sieve. Any new core chronology that fell within ±3 thousand years of the global stack of oxygen isotope records—known as the LR04 stack, compiled by Lisiecki and Maureen Raymo in 2005—was deemed compliant. Cores outside that band were flagged for re-analysis. The idea was to create a standard that reviewers and program managers could use to evaluate grant proposals and publications.

But the sieve had unintended consequences. Reviewers began to favor studies that used the LR04 stack as a tuning target, because it guaranteed that the resulting age model would pass the compliance check. Grant applications that proposed alternative tuning targets—for example, using regional instead of global stacks—were more likely to be rejected or asked for revision. A 2021 study by the budget office's internal evaluation unit found that the proportion of funded proposals using non-LR04 targets dropped from roughly 30% in 2014 to about 10% by 2019. One lab director, who spoke on condition of anonymity because his lab's funding was under review at the time, told a reporter for Science in 2020: “You don't get funded to disagree. If you propose a chronology that doesn't match the stack, reviewers assume you made a mistake.”

Publication pressure reinforced the pattern. Journals in paleoclimatology often require authors to state that their age models are consistent with the global stack. A 2021 survey of 150 papers published between 2015 and 2020 found that over 80% used the LR04 stack as either a primary or secondary tuning target. The sieve standard, originally intended to filter out obvious errors, had become a conformity sieve that discouraged exploration of alternative—and potentially more accurate—chronologies.

The budget office's own internal review, obtained through a Freedom of Information request, acknowledged that the tolerance threshold had “inadvertently reduced the diversity of age models in the literature.” But the office defended the sieve as a necessary compromise between methodological rigor and the practical need for consistency across studies.

The Sieve That Let Some Signals Through

The sieve standard was not just a bureaucratic tool; it actively shaped what counts as a valid climate signal. Consider a core from the Southern Ocean that showed a 4,000-year offset in the timing of the last deglaciation relative to the LR04 stack. Under the old rules, that offset would have been interpreted as a real regional difference—perhaps driven by local changes in sea ice or ocean circulation. Under the sieve standard, the core was re-analyzed until its age model fell within the ±3 kyr band, effectively erasing the apparent regional signal.

Critics, including paleoceanographer Peter Huybers of Harvard University, argued that the sieve standard was “a way to hide real variability behind a procedural rule.” Huybers pointed out that the LR04 stack itself is an average of many cores, and that individual cores can legitimately deviate from the stack by several thousand years due to local sedimentation effects. Forcing every core to match the stack, he said, would smooth out the very regional signals that paleoclimatologists need to understand past climate dynamics.

Supporters of the sieve countered that without a standard, the field would be awash in mutually incompatible chronologies, making it impossible to compare results across studies. “We have to have some baseline,” said a program director at the budget office who helped design the threshold. “Otherwise every paper becomes a special case and we lose the ability to synthesize.” The debate mirrored a similar tension in other fields, such as radiocarbon calibration, where a consensus curve is used despite known regional offsets.

Lisiecki's own position has evolved. In a 2023 interview, she said that the sieve standard was “a useful first step, but not a final solution.” Her team now advocates for what they call “multi-model chronologies,” in which multiple age models are produced for the same core using different tuning targets, and the spread of those models is reported as a measure of uncertainty. The approach is already being adopted by a handful of labs, including the one that produced the original Pacific core.

What the Divergence Reveals About Climate Science

Regional ocean circulation can shift sedimentation rates dramatically. In the Pacific, for example, the carbonate compensation depth varies by hundreds of meters across the basin, causing some cores to accumulate faster or slower than the global average. Local effects can mimic or mask global signals: a pulse of meltwater from Antarctica, for instance, can alter the oxygen isotope composition of seawater independently of ice volume. These local effects are not noise; they are the data we need to understand how different parts of the climate system respond to global forcing.

The sieve standard may have also discouraged methodological innovation. Labs that developed new tuning algorithms—for example, using Bayesian statistics to incorporate uncertainty—found that their age models often fell outside the ±3 kyr band, not because they were wrong, but because they incorporated different assumptions about sedimentation rate variability. These labs were less likely to receive funding, as grant reviewers applied rigid sample-size rules that favored conventional approaches.

Lisiecki's team now advocates for publishing raw age-depth pairs alongside the final age model, so that other researchers can test alternative tuning assumptions. Some labs already share their code for Bayesian age modeling, and a few journals have begun to require uncertainty envelopes on age-depth plots. But the cultural shift is slow. “The field is used to having one number for the age of a layer,” Lisiecki said. “It's hard to convince people that the uncertainty is part of the data.”

Tighter Sieves or Open-Access Chronologies?

The budget office's experiment with the sieve standard raises a deeper question: should paleoclimatology aim for tighter sieves that enforce consistency, or for open-access chronologies that embrace variability? The answer may depend on the question being asked. For global-scale studies of ice-volume changes, a consistent stack like LR04 is essential. But for regional studies of ocean circulation or carbon cycle dynamics, a single stack may obscure the very signals that matter.

One proposal, championed by a group of early-career researchers, is to create a public database of age-depth pairs and tuning parameters for every published core. Such a database would allow anyone to reconstruct an age model using their own assumptions, and to compare their results with others. The same principle has worked in microscopy, where sharing raw instrument data resolved discrepancies between labs. In paleoclimatology, the database would be a resource for meta-analyses and for testing the robustness of claims about the timing of past events.

But open-access chronologies come with risks. Without a standard, reviewers might be overwhelmed by multiple possible age models for the same core, and the literature could become fragmented. The budget office, for its part, has signaled that it will not renew the sieve standard after 2025. Instead, it plans to fund a series of workshops to develop community guidelines for reporting age-model uncertainty. The goal, according to a program director, is to move from a single threshold to a “spectrum of acceptable practices” that balance consistency with flexibility.

The real test, however, is whether the field can handle more disagreement. For decades, paleoclimatology has presented a unified story of glacial-interglacial cycles driven by orbital forcing. That story is largely correct, but it is built on a foundation of age models that may be more uncertain than most researchers realize. The two cores that started this story—one Pacific, one Atlantic—are now being re-dated using multiple tuning targets. Their new age models do not agree perfectly, but the spread of possible ages is reported openly. Whether that openness will lead to a more robust understanding of past climate, or simply to more confusion, remains an open question.

Recommend Posts
Science

One Desk Drawer Holding Two Magnetometer Calibration Constants

By Karim Osman/Jul 18, 2026

A 0.7% offset between two calibration constants for the same magnetometer sat unnoticed for years. How funding gaps and career incentives buried a discrepancy that could reshape exoplanet magnetic field estimates.
Science

A Single fMRI Slice-Timing Parameter Split Two Lab Anxiety Studies

By Karim Osman/Jul 18, 2026

Two labs studied anxiety with fMRI and got opposite results. The only difference was a slice-timing correction parameter in SPM. A replication audit reveals how a tiny default split the field.
Science

A Single Budget Overhead Cap Splits Two Climate Code Reproducibility Analyses

By Renu Shah/Jul 18, 2026

How a 15% NSF overhead cap on subawards forced two climate modeling teams down diverging paths—one verified, one retracted—over a difference of roughly $8,000 in compute funding.
Science

One Sieve Mesh Size Reassigned Two Hundred Polymer Viscosity Measurements

By Alice Chen/Jul 18, 2026

A single sieve mesh size change in a polymer lab reassigned 200 viscosity measurements. How procedural choices in materials science produce data that can shift 15–30%.
Science

A Grant Reviewer's Catalyst Purity Clause Scuttled Two Synthesis Labs

By Alice Chen/Jul 18, 2026

How a single clause demanding trace metal purity in catalysts derailed two synthesis labs, costing 18 months of work and sparking debate over funding gatekeeping.
Science

Twenty-Seven Crystal Growth Runs Traced One Superconductivity Reproducibility Gap

By Alice Chen/Jul 18, 2026

How a new instrument with real-time oxygen monitoring traced a reproducibility gap in superconducting crystal growth, and what it means for funding structures.
Science

A Budget Office’s Sieve Standard Split Two Ocean Sediment Core Chronologies

By Karim Osman/Jul 18, 2026

How a budget office's tolerance threshold for sediment core age models exposed hidden assumptions in paleoclimate dating, sparking debate over whether the field's consensus standard filtered out legitimate variability.
Science

Seven-Line Vignette Bias Shrank One Behavioral Economics Replication

By Alice Chen/Jul 18, 2026

A seven-line vignette in a 2012 behavioral economics study may have contributed to its failure to replicate. Weak treatments, small samples, and flexibility in analysis all played a role.
Science

A Single Lab Ventilation Glitch Altered Two Mouse Learning Curves

By Renu Shah/Jul 18, 2026

A lab HVAC failure skewed mouse memory data, revealing how cage microclimate masks genotype effects and threatens reproducibility in behavioral neuroscience.
Science

Twenty Psych Lab Equipment Budgets Forced One Replication Protocol Off a Second Registry

By Karim Osman/Jul 18, 2026

How equipment costs forced a replication protocol off PsychFileDrawer. The Many Eyes project aimed to replicate 20 studies; only 14 finished. Budgets, not theory, were the bottleneck.
Science

Ten Amphipod Species Disappeared from One Corrected Sediment Core Chronology

By Renu Shah/Jul 18, 2026

A revised chronology of a Lake Greifensee sediment core reveals ten amphipod species vanished abruptly, not gradually. The finding underscores how dating precision can flip paleoclimate narratives.
Science

A Single Tide Gauge Rental Fee Shifted Two Sea Level Acceleration Curves

By Alice Chen/Jul 18, 2026

How a $15,000 annual rental fee for two tide gauges, lost when a grant renewal failed, introduced a hidden discontinuity in Pacific sea level acceleration curves—and what it reveals about the fragile economics of long-term climate observation.
Science

One Climate Code Branch Forced Two Ocean Models Onto Different Turbulence Closures

By Alice Chen/Jul 18, 2026

A code fork around 2005 split two major ocean models onto different turbulence closure schemes. The choice of closure amplifies over decades, affecting hindcasts and reproducibility.
Science

A Gut Microbe Enzyme Rate Shifts Two Lab Mouse Anxiety Assays

By Alice Chen/Jul 18, 2026

A single gut bacterial enzyme produces contradictory results in two standard mouse anxiety tests, raising questions about how we measure anxiety-like behavior in rodents.
Science

Sixteen Bat Night-Roost Surveys Funded One Statistical Power Calculation

By Karim Osman/Jul 18, 2026

A single power calculation cost $4,800, funded by sixteen bat roost surveys. This article examines the hidden trade-offs in ecology research funding between data collection and statistical inference.
Science

A Single Photometric Calibration Star Divided Two Exoplanet Atmosphere Spectra

By Karim Osman/Jul 18, 2026

How one calibration star enabled two exoplanet atmosphere spectra, revealing water, sodium, and haze. A methodology piece on the hidden role of stellar templates.
Science

A Bureau of Land Management Drilling Fee Split Two Seismic Hazard Forecasts

By Karim Osman/Jul 18, 2026

A BLM drilling fee in Oklahoma triggered a split between two seismic hazard models—one from USGS, one industry-funded. The disagreement delayed a permit and exposed deeper tensions in earthquake risk assessment.
Science

Five Kaleidoscope Phase Plates Reveal One Quantum Optics Measurement Angle

By Alice Chen/Jul 18, 2026

A controversial quantum optics method uses five kaleidoscope phase plates to measure a single angle. Replication attempts reveal hidden parameters and procedural choices that split the field.
Science

A Flat Grant Overhead Rate Merged Two fMRI Anxiety Protocols

By Karim Osman/Jul 18, 2026

How a flat institutional overhead rate forced two anxiety fMRI studies into one hybrid protocol, diluting specificity and raising questions about funding incentives in neuroscience.
Science

A Sieve Mesh Size Swapped Two Ocean Circulation Records

By Renu Shah/Jul 18, 2026

Two studies analyzing the same sediment core reached opposite conclusions about Atlantic Ocean circulation. The culprit: a 63 µm versus 150 µm sieve mesh that selected different foraminifera size fractions.