Play X-ray spectra
The fit that can't tell
A neural network can fit an X-ray spectrum in milliseconds, so in my preprint I tried to catch one being wrong. A strong extra emission line got caught. A 3% error in the detector's energy scale got past all three checks I had, because the fit moves the source parameters to cover for it and the spectrum still looks fine. This toy does the same thing to one simulated source, live, with a plain likelihood fit standing in for the network.
Loading the detector response and fitting the first spectrum.
Fit terrain
how much worse the fit gets, over photon index and power-law norm- best fit
- truth
- best fit with no gain error
- the gain's push
- path as the gain moved
Posterior cloud
1,400 draws over Γ, norm and NH, Gaussian approximation- now
- no gain error
- truth
- cloud centre
- gain marginalised
Detector counts
XMM-Newton EPIC-pn channelsThe three checks
amber at p < 0.05, red at p < 0.01-
Replay test
Simulate 400 spectra from the fit. Does the real one look like them?
… -
Reweighting test
Score a rough first guess against the exact likelihood. If only a few draws survive, something is off.
… -
Evidence test
Could any setting of the model have made this spectrum?
…
A catalogue of 1,000
random sources from the paper's prior, each fitted with and without the gain error, same photonsEach source gets its own photons, fitted twice: once as recorded and once with the gain error. The histogram is the shift in fitted Γ. It takes up to half a minute.
What you're seeing
The spectrum goes through the real XMM-Newton EPIC-pn response I used in the paper, with the paper's source model: an absorbed power law plus a blackbody, five free parameters and no background. The gain slider stretches the energy scale the source is read on, so at 3% every feature lands 3% lower in energy than it should. The fit assumes the detector is fine and gets redone from scratch every time you move anything. Each channel gets Poisson counts, and the same random numbers are kept as you move the slider, so the only thing that changes is the gain. Switch photon noise off and the data become the exact expected counts. The fit then sits on the truth at zero gain, but the checks have nothing to judge, so they stand down.
It hides because a power law looks exactly the same after you stretch the energy axis, only its normalisation changes, and a stretched blackbody is just a slightly cooler blackbody. Nearly all of the gain error gets soaked up by the source parameters. What's left sits in the absorption edges, which don't move with the source, and that's where the photon index gets pushed. The terrain shows this: at every point it is the best fit you can get with Γ and the power-law norm pinned there, and the height is √ΔC, so each glowing contour is about one more sigma. As the gain grows the valley keeps its depth and the ball slides along it, far less than one contour at 3%, which is why the close-up in the corner exists. A strong iron line does the opposite: the whole fit gets worse, and the terrain turns red. For the default source at 10,000 counts, even a 10% gain error makes the best fit worse by about 1 in the Cash statistic, which noise does on its own.
The three lights are simplified versions of the three schemes in the paper. The replay test is the posterior-predictive check, a chi-square and a cumulative-counts comparison against 400 replayed spectra. The reweighting test follows the importance-sampling check, which reweights the network's posterior against the exact Poisson likelihood. The toy has no network, so its rough first guess is a fit to the spectrum squeezed into 20 broad bands (the network also squeezes each spectrum into 20 numbers), and the light compares how many draws survive with 100 clean spectra of the same source. The evidence test stands in for nested sampling's evidence, compared with clean spectra at the same counts. Here it's the best-fit Cash statistic against what a clean spectrum with those counts should give. On 150 clean noisy spectra of the default source each light went red 0 to 2% of the time, and the 3% gain didn't change that.
For one source the push is a small slice of the error bar, which is why nothing complains. The catalogue shows the other half. In the paper the network's photon index came out biased by +0.018 ± 0.006 at 3% and about 1,000 counts, and +0.019 at 10,000 counts, which is only 0.03 to 0.11σ per source but goes the same way for the whole sample.
The marginalise switch refits with the gain as a sixth unknown, flat between 0.95 and 1.05, like the retrained network in the paper. In the toy the gain usually comes back nearly as wide as its prior, the error on Γ barely moves and the push stays. The paper saw the same thing: the retrained network gave the prior back on the gain, its σ(Γ) changed by a factor of 0.989 to 0.996, and the bias was +0.020 ± 0.006. Exact nested sampling did get some gain information out of three bright spectra, and the network caught at most 0.163 of it. So marginalising doesn't remove the bias. It's still the defence the paper ends on, because it costs nothing when you fit and none of the checks warns you anyway. Past 5% the true gain is outside that prior, and marginalising has no way of knowing.
What this toy leaves out
The network. Every fit here is maximum likelihood on the Poisson statistic, so the toy behaves more like the paper's exact reference than like the network, and the shift it reports is the shift of the best fit, where the paper reports the shift of the network's posterior mean. The cloud is a Gaussian approximation around that fit, cut to the prior box. It's not the network's posterior either.
The response is thinned to keep the page small: pairs of input energy bins merged and the smallest matrix entries dropped. It matches the paper's simulator (jaxspec) to within 2% per channel, and 3% in the channel holding the iron line.
The terrain is computed on a 15 by 15 grid and smoothed in between, and its axes widen when the fit would otherwise leave the map. The default noise pattern is one I picked because its zero-gain fit lands just above the truth (Γ = 2.002), so the gain visibly pushes it further away, and its shift at 3% (+0.015) sits inside the range the catalogue gives; press New noise for others. The checks are cut-down versions, one source is fitted at a time, and the gain is a slope with no offset, same as in the paper.
With the noise taken out, the toy's shift at 3% doesn't depend on exposure at all. Over 1,000 random sources from the prior its median is +0.021 and 93% of sources move the positive way. The mean is only +0.011 ± 0.002, because about 3% of sources, mostly ones where the blackbody outshines the power law so Γ is barely pinned down, get shoved by 0.1 or more, nearly all of them negative. So the number the catalogue reports weights each source by 1/σ² of its fitted Γ, the precision weighting the paper also uses for its second bias test, and those loose sources count for little. With noise, ten catalogues of 1,000 gave weighted means of +0.014 to +0.018 at about 1,000 counts and +0.014 to +0.016 at 10,000, each at least 8σ from zero. The unweighted means spread wider, +0.010 to +0.025 and +0.013 to +0.021, but all ten were still more than 2σ above zero at both levels. Both sit a little under the paper's +0.018 and +0.019. It doesn't have to match, since one is a best fit and the other a network's posterior mean.
The longer write-up of the paper
The preprint, arXiv:2606.17098
References
- Akbari, K. 2026, arXiv:2606.17098 (preprint). Misspecification in amortized X-ray spectral inference: a detection benchmark, a gain shift three schemes cannot see, and marginalizing it out.
- Barret, D., & Dupourqué, S. 2024, A&A, 686, A133. Simulation-based inference with neural posterior estimation applied to X-ray spectral fitting (source of the priors).
- Cash, W. 1979, ApJ, 228, 939. Parameter estimation in astronomy through application of the likelihood ratio.
- Dupourqué, S., Barret, D., Diez, C. M., Guillot, S., & Quintin, E. 2024, A&A, 690, A317. jaxspec: a fast and robust Python library for X-ray spectral fitting (the simulator the response comes from).
- Wilms, J., Allen, A., & McCray, R. 2000, ApJ, 542, 914. On the absorption of X-rays in the interstellar medium (the tbabs cross-sections).