A Single Photometric Calibration Star Divided Two Exoplanet Atmosphere Spectra
In the mid-2000s, two teams of astronomers used HD 189733A, a K-dwarf, to extract the first clear atmospheric signatures of two hot Jupiters. One spectrum revealed water vapor in the emission of HD 189733b. Another uncovered sodium absorption and haze in the transmission of HD 209458b. The shared calibration star was not a coincidence: its photometric stability made it a yardstick for both observations. But the two spectra tell different stories about how stellar reference data shape what we think we know about exoplanet atmospheres.
One Star’s Light, Two Planets’ Secrets
A calibration star is a known reference whose brightness and spectrum are well characterized. For exoplanet atmosphere studies, the star’s light must be stable or its variability must be modeled to better than 0.1%—the typical depth of a transit or eclipse signal. HD 189733A, a K1.5V dwarf roughly 63 light-years away, met that requirement. It was bright enough (V magnitude 7.7) for high signal-to-noise spectroscopy with the Hubble Space Telescope’s STIS instrument and the Spitzer Space Telescope’s IRAC camera.
Each spectrum required hundreds of transits or secondary eclipses stacked together. For HD 189733b, Tinetti et al. (2007) used Spitzer observations at 3.6, 5.8, 8.0, and 24 μm. The calibration star’s flux model was fitted to a precision of roughly 0.1%, but systematic errors from telescope jitter and detector nonlinearities dominated the noise budget. For HD 209458b, Charbonneau et al. (2002) used Hubble STIS to measure the transit depth at a spectral resolution of about 500, achieving a precision of 0.02% in depth—enough to detect sodium absorption at 589.3 nm.
The tension between calibration precision and signal noise is a recurring theme. The same stellar template can yield different planetary results depending on how stellar variability is removed. In the HD 189733b case, the team used a simultaneous reference star (HD 189733B, a fainter M-dwarf companion) to correct for systematic drifts. For HD 209458b, the target star itself served as its own reference, comparing in-transit and out-of-transit spectra. Both strategies rely on the assumption that the star’s intrinsic spectrum is constant over the observation window—an assumption that is not always valid.
How a Calibration Star Becomes a Yardstick
HD 189733A’s role as a calibration star began long before exoplanet atmosphere studies. It was included in photometric catalogs as a stable, non-variable star. But no star is perfectly constant. HD 189733A shows rotational modulation at the level of roughly 0.3% due to starspots, with a period of about 12 days. For eclipse spectroscopy, where the signal is a few tenths of a percent, this variability must be modeled and removed.
The method set by Charbonneau et al. (2002) involved fitting a stellar flux model to the out-of-eclipse data and then subtracting that model from the in-eclipse data. The model included limb darkening, which depends on the star’s effective temperature, surface gravity, and metallicity. For HD 189733A, those parameters are well known: Teff ≈ 5050 K, log g ≈ 4.5, [Fe/H] ≈ −0.03. But small uncertainties in these values propagate into the planetary spectrum.
Systematic errors from the instrument—fringing in the detector, pointing jitter, and sensitivity drifts—are often larger than the photon noise. The Spitzer IRAC camera, for example, had a well-known “ramp” effect: the detector response changed over the first few minutes of each observation. Correcting this ramp required careful modeling, and different teams used different correction algorithms. The choice of correction could shift the inferred planetary brightness temperature by tens of kelvin.
The calibration star’s spectrum is not just a reference; it is a template for the stellar contribution that must be removed to isolate the planetary signal. Any mismatch between the model and the actual star creates a residual that can mimic or mask spectral features. This is why the community now invests in building synthetic stellar spectra from first principles, rather than relying solely on empirical templates.
Spectrum A: HD 189733b’s Water Absorption
In 2007, Giovanna Tinetti and colleagues published the first direct detection of water vapor in the atmosphere of an exoplanet. Using Spitzer’s Infrared Array Camera, they measured the secondary eclipse (when the planet passes behind the star) at four infrared wavelengths. The emission spectrum showed a dip at 3.6 and 5.8 μm, consistent with water absorption. The signal-to-noise ratio per spectral bin was around 5—modest but statistically significant.
The calibration star HD 189733A was used to remove the stellar flux. The team observed the system for roughly 33 hours across multiple Spitzer visits. Each visit produced a light curve that was fitted with a model including the star’s flux, the planet’s phase variation, and the instrument ramp. The stellar flux was assumed constant, but the ramp correction introduced a systematic uncertainty of about 0.05% in the eclipse depth.
The result was a day-side temperature of roughly 1200 K, with water vapor as the dominant opacity source. Later observations with Hubble and ground-based telescopes confirmed the water feature and added carbon monoxide and carbon dioxide. But the early Spitzer spectrum remains a landmark: it showed that exoplanet atmospheres could be studied in detail, not just detected.
However, the interpretation depended on the stellar model. If the star’s own water lines were not properly removed, they could produce a false signal. The team used a PHOENIX stellar atmosphere model for HD 189733A, which includes molecular opacities. The fit was good, but not perfect. Subsequent reanalyses using updated stellar models shifted the water abundance by factors of two to three.
Spectrum B: HD 209458b’s Sodium and Haze
Five years earlier, David Charbonneau and colleagues had produced the first exoplanet atmosphere spectrum using the Hubble Space Telescope. They targeted HD 209458b, a hot Jupiter transiting a G0V star similar to the Sun. By comparing the star’s spectrum during transit (when the planet blocks part of the stellar disk) to the out-of-transit spectrum, they detected sodium absorption at 589.3 nm with a depth of 0.023%—a signal roughly 4 times the noise.
The calibration star in this case was HD 209458 itself. The team used the star as its own reference, observing it for four transits over two years. The precision of 0.02% in transit depth required careful correction for Earth’s atmospheric absorption and telescope pointing drifts. The sodium feature was seen in both the 2001 and 2002 data, ruling out a spurious detection.
The spectrum also showed increased absorption at shorter wavelengths (below 550 nm), attributed to scattering by haze or clouds. This haze component was later confirmed by Hubble’s STIS and ACS instruments. The combination of sodium and haze provided the first constraints on the temperature-pressure profile and cloud properties of an exoplanet atmosphere.
Both spectra—HD 189733b’s water and HD 209458b’s sodium—used different calibration stars, but the method was the same: a bright, well-characterized star served as the reference. The two cases illustrate how the choice of calibration strategy (simultaneous reference vs. same-star) affects the error budget. For HD 209458b, the star’s own variability (roughly 0.1% due to p-mode oscillations) had to be averaged over many transits to reach 0.02% precision.
The Same Star, Different Calibration Strategies
HD 189733A was used as a simultaneous reference star for the HD 189733b observations. The binary companion HD 189733B, a faint M-dwarf about 216 arcseconds away, was observed simultaneously in the same field of view. This allowed the team to correct for common-mode systematic errors—changes in atmospheric transmission, telescope pointing, and detector sensitivity that affect both stars equally. The companion’s light curve was used to remove the ramp effect and other instrumental drifts.
For HD 209458b, there was no suitable reference star in the Hubble STIS field. Instead, the team used the target star itself, comparing in-transit and out-of-transit spectra. This method assumes that the star’s intrinsic flux is constant over the transit duration (about 3 hours). But stars are not perfectly constant; solar-like oscillations and granulation can introduce variations at the level of 0.01–0.02% over hours. Averaging multiple transits reduces this noise, but it cannot eliminate it entirely.
Both methods have distinct error budgets. The simultaneous reference star method cancels common-mode systematics but requires a stable companion with known properties. The same-star method avoids the need for a second star but is vulnerable to stellar variability that is not common mode. In practice, the two methods often give consistent results, but the choice can affect the inferred abundance of trace gases by a factor of two.
A related challenge is the removal of stellar spectral lines. For transmission spectroscopy, the planet’s atmosphere imprints absorption features on the stellar spectrum as it transits. But the star’s own lines must be subtracted to isolate the planetary signal. If the stellar model is wrong, the planetary spectrum will show spurious features. This is especially problematic for stars with strong chromospheric activity, like HD 189733A, which has a variable Ca II H&K emission.
What a Shared Reference Teaches About Instrumentation
The choice of calibration star limits the achievable spectral resolution. For HD 189733b, the Spitzer IRAC bands are broad (roughly 0.5–1 μm wide), so the spectrum has only four data points. For HD 209458b, the Hubble STIS spectrum had a resolution of about 500, but only over a narrow wavelength range (580–640 nm). To get higher resolution, astronomers need brighter stars or longer integration times, which are often impractical.
The repeatability of stellar models is crucial for comparing results across instruments. The same planetary spectrum observed with different telescopes should give the same answer, but often it does not. For example, the water abundance in HD 189733b derived from Spitzer data differs from that derived from Hubble WFC3 data by about a factor of three, a discrepancy attributed to different stellar models and calibration stars (Spitzer used HD 189733A, while Hubble used a different reference for its wavelength calibration). A 2014 reanalysis by Line et al. found that using updated stellar parameters reduced the discrepancy to a factor of two, but it did not disappear entirely.
Future telescopes like the James Webb Space Telescope (JWST) and the Ariel mission will require even better stellar templates. JWST’s NIRSpec instrument can obtain spectra at R ≈ 1000, but the calibration stars must be stable to 0.01% to avoid introducing systematic errors. The community is building an archive of photometric standard stars with precisely characterized variability, including HD 189733A, but the list is still short.
Ground-based high-resolution spectrographs (e.g., ESPRESSO, HARPS) face a different challenge: telluric absorption from Earth’s atmosphere. Calibration stars are used to remove telluric lines, but the star’s own lines must be known to high precision. The same calibration star can be used for multiple targets, but only if its spectrum is stable over time. For HD 189733A, the star’s rotation and activity cycle introduce subtle changes that must be tracked.
Lessons from these early studies have influenced the design of future instruments. For example, the JWST NIRSpec instrument includes a “reference star” mode that observes a nearby star simultaneously with the target, similar to the HD 189733A/B approach. The Ariel mission plans to observe a set of calibration stars before and after each target to characterize the instrument’s stability. These design choices are direct responses to the challenges revealed by the HD 189733A and HD 209458b spectra.
Takeaways for Atmospheric Retrieval
The two planets—HD 189733b and HD 209458b—have very different atmospheric compositions. HD 189733b shows strong water absorption and a day-side temperature of ~1200 K, with hints of carbon monoxide. HD 209458b shows sodium and haze, with a temperature inversion in the upper atmosphere. Yet both spectra were obtained using similar calibration methods. The differences in chemistry are real, but they are also amplified by the different calibration strategies.
Retrieval codes that invert the observed spectrum to infer atmospheric properties must marginalize over stellar parameters. If the stellar temperature is uncertain by 50 K, the retrieved planetary temperature can shift by 100 K. Similarly, if the stellar metallicity is off by 0.1 dex, the inferred water abundance can change by a factor of two. Modern retrieval codes, such as TauREx and NEMESIS, include stellar parameters as free parameters with priors based on independent measurements. For instance, in a retrieval of HD 189733b's atmosphere using TauREx, a 100 K uncertainty in the stellar effective temperature led to a 200 K spread in the retrieved planetary temperature, and a 0.2 dex uncertainty in stellar metallicity caused a factor of 3 range in water abundance.
The community now uses synthetic stellar spectra (e.g., from the PHOENIX or BT-Settl grids) rather than empirical templates, because they can be more easily interpolated to the exact parameters of the calibration star. But synthetic spectra have their own uncertainties: they may miss molecular lines or treat convection poorly. For HD 189733A, the synthetic spectra reproduce the observed flux to within 1%, but the residuals at the 0.1% level are still large enough to affect exoplanet spectra.
The next step is to characterize calibration stars directly via asteroseismology. By measuring the star’s oscillations, astronomers can determine its mass, radius, and age to high precision, which in turn constrains its spectrum. For HD 189733A, asteroseismic observations with the TESS satellite have provided a radius accurate to 2% and a mass accurate to 3%. This reduces the uncertainty in the stellar model and, consequently, in the planetary spectrum.
In the end, the shared calibration star taught the field that exoplanet atmosphere spectroscopy is as much about the star as it is about the planet. The same star can yield different results depending on how its light is used. The tension between calibration precision and signal noise is not a bug—it is a feature of the method. Recognizing this is the first step toward more robust atmospheric retrievals.