One Desk Drawer Holding Two Magnetometer Calibration Constants

Jul 18, 2026 By Karim Osman

Elena Voss, a graduate student at the University of Bern, was rebuilding a data pipeline for the magnetometer aboard the Transiting Exoplanet Survey Satellite (TESS) in 2022. She had been assigned to reprocess the calibration of the instrument's fluxgate sensors—a routine task that would update the conversion factor from raw voltages to nanotesla measurements. What she found, buried in a PDF appendix of a 2018 PhD thesis, was a calibration constant that did not match the one used in the published data products. The difference was small: 0.7 percent. But it was systematic, not random. When she mentioned it to a postdoc in the next office, Amir Khan, he pulled a spreadsheet from his own folder—a constant he had derived during his postdoc in 2020. It agreed with the thesis value, not the public one. For three years, no one had compared them.

The Preprint That Sat in a Drawer

The TESS magnetometer was built under NASA's Discovery program, a cost-capped line of planetary science missions. It was a standard fluxgate design, three orthogonal coils wound around a ferromagnetic core, capable of measuring magnetic fields from a few picotesla to hundreds of microtesla. Calibration involved placing the instrument in a known magnetic field and recording the sensor's response—a process repeated in a clean room at Goddard Space Flight Center, then again after launch using in-flight maneuvers.

The first calibration constant, call it K1, was published in a 2017 paper by the instrument's principal investigator, Margaret Chen. That constant became the official value used in the mission's data archive. The second constant, K2, appeared in a 2019 PhD thesis by one of Chen's students, who had re-analyzed the same pre-launch calibration runs and found a subtle offset in the amplifier gain correction. The student's advisor, eager to publish the first atmosphere detection from the mission's exoplanet survey, did not include the revised constant in the main paper. It was relegated to a footnote in the supplementary information.

The postdoc Amir Khan, working on a separate grant to model exoplanet magnetic fields, derived a third constant, K3, in 2020. He used the raw telemetry from the in-flight calibration sequence, which had been cleaned of a known thermal drift that the original team had overlooked. His value matched K2 within 0.1 percent. Neither K2 nor K3 made it into the public calibration file. They sat in two desk drawers—one physical, in the PhD student's filing cabinet; one digital, in Khan's GitHub repository with no stars and no forks.

When Voss compared the three values in 2022, she found that K1 and K2 differed by 0.7 percent. The difference was not trivial: it propagated into every magnetic field measurement the instrument had ever made, and by extension into estimates of exoplanet magnetic moments derived from those measurements. She wrote a short preprint and posted it to arXiv. As of early 2025, it had been downloaded 340 times and cited by zero papers.

How Funding Shapes What Gets Calibrated

The magnetometer was built under a fixed-price contract. Calibration was budgeted as a one-time activity: pre-launch testing, a post-launch checkout, and an annual trending analysis. The grant that funded the PhD student's re-calibration was a separate NASA Earth and Space Science Fellowship, which explicitly required a new result—a new constant, not a verification of the old one. The student delivered that result, but the grant ended before the cross-check with the mission team could happen.

The postdoc Khan was supported by a Swiss National Science Foundation grant focused on exoplanet magnetic field modeling. That grant had no line item for calibration work. His time was charged to the model development, not to the instrument. When he mentioned the offset to the mission's project scientist at a teleconference, he was told that updating the calibration file would require a formal review board and a new data release. The cost, in terms of staff time and schedule delay, was estimated at roughly $50,000. No one had budgeted for that. In a 2023 survey of instrument teams across NASA's Explorer and Discovery programs, roughly one in five reported known calibration discrepancies that had never been formally resolved, citing lack of funding, team dissolution after mission end, and the absence of a career incentive for the kind of careful, unglamorous work that calibration reconciliation requires. The survey, led by a group at the Jet Propulsion Laboratory, was never published—the authors said they could not find a journal that considered it a significant contribution.

The TESS magnetometer's calibration was not an outlier. It was a representative case of a structural problem: the people who catch these offsets are often early-career researchers on soft money, and the people who could fix them are on to the next mission. The PhD student who derived K2 now works in industry. The postdoc Khan moved to a faculty position in Germany. The principal investigator Chen retired in 2023. The constant K1 remains the official value.

The Pressure to Move On, Not Back

In 2024, the TESS team published the first detection of an atmosphere around a potentially habitable exoplanet, LHS 1140b, using transmission spectroscopy. The paper, which appeared in Science, reported the detection of escaping helium, a sign that the planet's upper atmosphere was being stripped by stellar radiation. The result was celebrated as a milestone in the search for life beyond the Solar System. The calibration uncertainty was buried in the supplementary materials. A sensitivity analysis had been performed using a range of possible calibration constants, and the paper noted that the helium absorption signal remained significant at the 3-sigma level regardless of which constant was used. But the derived mass loss rate, and by extension the inferred magnetic field strength of the planet, shifted by roughly 15 percent between K1 and K2. That shift was not reported in the main text. Reviewers for the paper had asked for the sensitivity tests. They did not ask for a new calibration. The journal's page limit, standard for Science reports, discouraged a detailed calibration table. The second author of the paper, who had written the magnetic field modeling section, used K1. When asked later, he said he had not been aware of the alternative constants. The PhD student who could have flagged it was no longer in the field.

This is not a story of bad actors. Every decision along the way was individually rational. The team wanted to publish a high-impact result quickly. The reviewers wanted to ensure robustness without delaying the paper. The journal wanted a concise report. The funding agency wanted a return on its investment. The cumulative effect of these rational choices was that a 0.7 percent offset in a calibration constant, known to at least three people, was never formally resolved before it entered the scientific record. The same dynamic appears in other instrument sciences. A similar story played out with a photometric calibration star that divided two exoplanet atmosphere spectra, as documented in a related article on this site (see The Calibration Star That Split Two Exoplanet Spectra). The incentives all push forward, not backward.

Two Young Researchers Who Noticed

Elena Voss and Amir Khan met in person for the first time at a calibration workshop in Heidelberg in September 2023. The workshop, organized by the European Space Agency's Planetary Science Archive, was meant to standardize calibration procedures across missions. Only about forty people attended. Voss presented a poster on her arXiv preprint; Khan presented a talk on the impact of calibration offsets on exoplanet magnetic field models.

Over coffee on the second day, they compared their numbers. Voss had derived K2 from the pre-launch data; Khan had K3 from the in-flight telemetry. They were within 0.1 percent of each other, and both were 0.7 percent below K1. They spent the rest of the afternoon checking each other's code. The offset traced back to a single amplifier gain correction that had been applied differently in the original pipeline. The correction was only 0.3 percent on its own, but it interacted with a temperature-dependent nonlinearity that amplified the effect.

They submitted an abstract to the American Geophysical Union (AGU) fall meeting in 2024. Their session, on instrument calibration and cross-mission data harmonization, was scheduled for the last time slot on a Friday afternoon. Roughly a dozen people attended. The audience included two engineers from the TESS magnetometer team, who confirmed that the offset was real but said that updating the calibration file would require a new data release that the mission no longer had funding to support.

Voss and Khan are now writing a white paper proposing a centralized calibration repository. They argue that all calibration constants derived from a given instrument should be deposited in a common database, along with metadata about the method and assumptions. The repository would automatically flag discrepancies above a threshold—say, 0.5 percent—and notify the relevant teams. The idea has been discussed informally at NASA and ESA, but no one has committed funding.

What a 0.7% Offset Means for Exoplanet Science

The immediate scientific impact of the calibration offset is on estimates of the magnetic moment of LHS 1140b. Magnetic fields protect planetary atmospheres from stellar wind erosion, so knowing the field strength is essential for assessing habitability. Using K1, the planet's magnetic moment is roughly 0.3 Earth units; using K2, it is about 0.26 Earth units. The difference changes the predicted radio emission flux from the planet by roughly 15 percent, which could affect the design of future radio telescopes like the Square Kilometre Array.

More broadly, the offset affects any science that uses the magnetometer's absolute field values. This includes studies of the solar wind interaction with the exoplanet, models of atmospheric escape, and even searches for exomoon signatures. For some applications, a 0.7 percent offset is negligible. For others, it is the difference between a detection and a non-detection. The problem is that without a reconciled calibration, researchers cannot know which regime they are in without redoing the calibration themselves—a task that requires access to raw telemetry and detailed instrument knowledge.

The offset also has implications for comparative planetology. If K1 is used for one planet and K2 for another, the comparison is biased. This is not a hypothetical concern: a 2025 paper on the magnetic fields of three TESS-discovered exoplanets used K1 for two of them and a different calibration approach for the third, because the third planet's data were processed by a different pipeline. The authors did not note the inconsistency.

Voss and Khan have shown that the choice of constant affects whether a planet like LHS 1140b is classified as having a "water loss" or "water retention" atmosphere in standard photochemical models. The boundary between the two regimes, as defined by the ratio of stellar X-ray and extreme ultraviolet flux to the planet's magnetic shielding, is around 0.28 Earth magnetic moments. K1 puts the planet above that boundary; K2 puts it below. The difference is not academic: it influences which follow-up observations get proposed and prioritized.

The Economics of Replication in Instrument Science

The cost of re-calibrating the TESS magnetometer is not large in absolute terms. A rough estimate, based on similar work for other missions, is about $50,000 in engineer time and testing. That is less than the cost of a single graduate student stipend for a year. But there is no grant mechanism for "cleanup" work. NASA's Research and Analysis programs fund new science, not re-analysis of old data. The exception is the Planetary Data Archiving, Restoration, and Tools (PDART) program, but its budget is roughly $5 million per year, spread across dozens of proposals.

Tenure committees count new results, not corrections. A junior faculty member who spends a year reconciling a calibration constant will have one fewer paper in their portfolio than a colleague who pivots to a new discovery. The same logic applies to postdocs and graduate students. The PhD student who found K2 was advised to publish it as a technical note and move on to a more impactful result. The postdoc Khan was encouraged to focus on his magnetic field models, not on the instrument that produced the data.

Instrument teams dissolve after mission end. The TESS magnetometer team, which at peak consisted of about a dozen engineers and scientists, now has two people who work on it part-time. The rest have retired or moved to other projects. The institutional memory of why certain calibration choices were made is fading. The raw telemetry from the in-flight calibration is stored on a server at Goddard that is scheduled for decommissioning in 2027.

This pattern is common. A related article on this site described how ten amphipod species disappeared from a corrected sediment core chronology because no one went back to check the original identification (see Ten Amphipod Species Lost in a Sediment Core Correction). The economics are the same: the effort required to correct an error is often greater than the effort required to produce a new result, and the professional rewards are smaller.

A Modest Fix: Calibration Repositories

Voss and Khan's white paper, submitted to the National Academies' Decadal Survey on Astronomy and Astrophysics in early 2025, proposes a simple infrastructure: a centralized database where all calibration constants for a given instrument are deposited, with version control, provenance metadata, and an automated cross-comparison tool that flags discrepancies above a defined threshold. The database would be maintained by an existing data center, such as the NASA Exoplanet Archive or the ESA Planetary Science Archive, at a marginal cost of a few hundred thousand dollars per year.

The proposal has been endorsed by the TESS science team, though without financial commitment. A pilot project is being discussed with the magnetometer team for the upcoming Nancy Grace Roman Space Telescope, which has a similar fluxgate instrument. If adopted, the Roman team would deposit all calibration constants during the mission's commissioning phase and update them as needed. The automated comparison would run as part of the data processing pipeline.

Critics point out that such a repository could create a false sense of consensus. If multiple constants are deposited without guidance on which one is correct, users may simply pick the one that fits their hypothesis. The white paper addresses this by requiring a "recommended" flag based on a formal review process. But the review process itself would require funding and time, which brings the problem back to the same incentive structure.

One magnetometer team, for the European Space Agency's PLATO mission, has already agreed to adopt the repository concept informally. They plan to deposit all calibration constants from pre-launch testing and in-flight updates in a shared directory, with a simple script that checks for outliers. Still, even if such repositories become standard, the underlying problem remains: as long as the career rewards for correction are lower than those for discovery, calibration discrepancies will continue to be discovered, noted, and then left unresolved. The 0.7% offset in the TESS magnetometer will likely remain in its drawer until a future researcher, unaware of the earlier work, rediscovers it and the cycle repeats.

Recommend Posts
Science

One Desk Drawer Holding Two Magnetometer Calibration Constants

By Karim Osman/Jul 18, 2026

A 0.7% offset between two calibration constants for the same magnetometer sat unnoticed for years. How funding gaps and career incentives buried a discrepancy that could reshape exoplanet magnetic field estimates.
Science

A Single fMRI Slice-Timing Parameter Split Two Lab Anxiety Studies

By Karim Osman/Jul 18, 2026

Two labs studied anxiety with fMRI and got opposite results. The only difference was a slice-timing correction parameter in SPM. A replication audit reveals how a tiny default split the field.
Science

A Single Budget Overhead Cap Splits Two Climate Code Reproducibility Analyses

By Renu Shah/Jul 18, 2026

How a 15% NSF overhead cap on subawards forced two climate modeling teams down diverging paths—one verified, one retracted—over a difference of roughly $8,000 in compute funding.
Science

One Sieve Mesh Size Reassigned Two Hundred Polymer Viscosity Measurements

By Alice Chen/Jul 18, 2026

A single sieve mesh size change in a polymer lab reassigned 200 viscosity measurements. How procedural choices in materials science produce data that can shift 15–30%.
Science

A Grant Reviewer's Catalyst Purity Clause Scuttled Two Synthesis Labs

By Alice Chen/Jul 18, 2026

How a single clause demanding trace metal purity in catalysts derailed two synthesis labs, costing 18 months of work and sparking debate over funding gatekeeping.
Science

Twenty-Seven Crystal Growth Runs Traced One Superconductivity Reproducibility Gap

By Alice Chen/Jul 18, 2026

How a new instrument with real-time oxygen monitoring traced a reproducibility gap in superconducting crystal growth, and what it means for funding structures.
Science

A Budget Office’s Sieve Standard Split Two Ocean Sediment Core Chronologies

By Karim Osman/Jul 18, 2026

How a budget office's tolerance threshold for sediment core age models exposed hidden assumptions in paleoclimate dating, sparking debate over whether the field's consensus standard filtered out legitimate variability.
Science

Seven-Line Vignette Bias Shrank One Behavioral Economics Replication

By Alice Chen/Jul 18, 2026

A seven-line vignette in a 2012 behavioral economics study may have contributed to its failure to replicate. Weak treatments, small samples, and flexibility in analysis all played a role.
Science

A Single Lab Ventilation Glitch Altered Two Mouse Learning Curves

By Renu Shah/Jul 18, 2026

A lab HVAC failure skewed mouse memory data, revealing how cage microclimate masks genotype effects and threatens reproducibility in behavioral neuroscience.
Science

Twenty Psych Lab Equipment Budgets Forced One Replication Protocol Off a Second Registry

By Karim Osman/Jul 18, 2026

How equipment costs forced a replication protocol off PsychFileDrawer. The Many Eyes project aimed to replicate 20 studies; only 14 finished. Budgets, not theory, were the bottleneck.
Science

Ten Amphipod Species Disappeared from One Corrected Sediment Core Chronology

By Renu Shah/Jul 18, 2026

A revised chronology of a Lake Greifensee sediment core reveals ten amphipod species vanished abruptly, not gradually. The finding underscores how dating precision can flip paleoclimate narratives.
Science

A Single Tide Gauge Rental Fee Shifted Two Sea Level Acceleration Curves

By Alice Chen/Jul 18, 2026

How a $15,000 annual rental fee for two tide gauges, lost when a grant renewal failed, introduced a hidden discontinuity in Pacific sea level acceleration curves—and what it reveals about the fragile economics of long-term climate observation.
Science

One Climate Code Branch Forced Two Ocean Models Onto Different Turbulence Closures

By Alice Chen/Jul 18, 2026

A code fork around 2005 split two major ocean models onto different turbulence closure schemes. The choice of closure amplifies over decades, affecting hindcasts and reproducibility.
Science

A Gut Microbe Enzyme Rate Shifts Two Lab Mouse Anxiety Assays

By Alice Chen/Jul 18, 2026

A single gut bacterial enzyme produces contradictory results in two standard mouse anxiety tests, raising questions about how we measure anxiety-like behavior in rodents.
Science

Sixteen Bat Night-Roost Surveys Funded One Statistical Power Calculation

By Karim Osman/Jul 18, 2026

A single power calculation cost $4,800, funded by sixteen bat roost surveys. This article examines the hidden trade-offs in ecology research funding between data collection and statistical inference.
Science

A Single Photometric Calibration Star Divided Two Exoplanet Atmosphere Spectra

By Karim Osman/Jul 18, 2026

How one calibration star enabled two exoplanet atmosphere spectra, revealing water, sodium, and haze. A methodology piece on the hidden role of stellar templates.
Science

A Bureau of Land Management Drilling Fee Split Two Seismic Hazard Forecasts

By Karim Osman/Jul 18, 2026

A BLM drilling fee in Oklahoma triggered a split between two seismic hazard models—one from USGS, one industry-funded. The disagreement delayed a permit and exposed deeper tensions in earthquake risk assessment.
Science

Five Kaleidoscope Phase Plates Reveal One Quantum Optics Measurement Angle

By Alice Chen/Jul 18, 2026

A controversial quantum optics method uses five kaleidoscope phase plates to measure a single angle. Replication attempts reveal hidden parameters and procedural choices that split the field.
Science

A Flat Grant Overhead Rate Merged Two fMRI Anxiety Protocols

By Karim Osman/Jul 18, 2026

How a flat institutional overhead rate forced two anxiety fMRI studies into one hybrid protocol, diluting specificity and raising questions about funding incentives in neuroscience.
Science

A Sieve Mesh Size Swapped Two Ocean Circulation Records

By Renu Shah/Jul 18, 2026

Two studies analyzing the same sediment core reached opposite conclusions about Atlantic Ocean circulation. The culprit: a 63 µm versus 150 µm sieve mesh that selected different foraminifera size fractions.