Can a fast X-ray fit tell when it's wrong?
Neural networks can fit an X-ray spectrum in milliseconds, which is great, as long as the data looks like what they trained on.
I simulated X-ray spectra with problems added in, like an extra emission line or a shifted detector energy scale, to see which ones the network would notice. The extra line was easy to spot. The shifted energy scale wasn't, because the fit just adjusts the source parameters to make up for it.
The energy scale shift got past all three checks, which is a little inconvenient for a method whose whole point is that you can trust it.
The longer story
The spectra are simulated, but they go through the real response files of XMM-Newton and NICER, so the detector side is as realistic as I could make it. And since I made the spectra, I know the right answer for every one of them.
I added four kinds of problems. An extra emission line, some absorbing gas partly covering the source, a different continuum shape, and a gain shift, which is when the detector's energy scale is slightly off.
An extra line is easy to catch because the model has no way to make a pile of extra counts at one energy. A gain shift stretches the whole spectrum a bit, and the fit handles that by moving the source parameters. So the fit still looks good and only the answer changes.
More photons made the line easier to catch and left the gain shift near chance at every count level.
I also checked whether the network's error bars were honest. If it says 68%, the true value should land inside that range 68% of the time. One of the trained networks was overconfident, its 90% intervals only caught the truth 76% of the time. Recalibrating brought that up to 88%, though that doesn't fix a wrong model.
Then I retrained it with the gain included as something it has to account for. That only helps when the spectrum actually carries information about the gain. When it doesn't, the gain just stays wherever the prior put it and the bias stays too. It's a small bias, 0.03 to 0.11 sigma, but it doesn't go away.
On a few brighter spectra, a slower exact method (nested sampling with the gain marginalized) pulled out gain information that the network mostly missed. So some of the problem is the data and some of it is the network.
For astronomers: setup, metrics, results
- Simulation
jaxspecthrough an XMM-Newton EPIC-pn response (102 grouped channels) and a NICER XTI response (969 channels, 0.3–10 keV). Source modeltbabs×(powerlaw+blackbody), 5 free parameters. No background component.- Estimator
- One neural spline flow per count level (
sbi), 50,000 simulations, CNN embedding to 20 summaries, 5 spline transforms - Count levels
- Median total counts ≈100, 1,000, 10,000
- Injected errors
- B1: 6.4 keV line (σ = 0.05 keV), 5×10−6 to 3×10−4 ph cm−2 s−1. B2: partial covering, fraction 0.9 to 0.3. B3: bremsstrahlung continuum, kT = 10 to 1.5 keV. B4: gain shift 0.5–3%, slope only.
- Detection
- AUC (0.5 = chance). B1: 0.76, 0.97, 0.97 at faint, medium, bright. Best per-spectrum check (D1 or D2): B2 up to 0.84, B3 up to 0.66. The population test D3 reaches 0.96 and 0.81 on those two, but it is not a per-spectrum score (Table 1). B4 over all 36 cells: 0.43–0.58, mean 0.50 (NICER mean 0.49). ESS stays null-consistent from 0.1 to 10% gain.
- Evidence
- Nested sampling, paired gain test at medium counts (n = 12): Δln Z = +0.33 ± 1.37 nats (p = 0.81). Line at medium counts: −67 nats [−90, −44].
- Calibration
- Bright EPIC-pn flow coverage deviation 0.114 (faint 0.014, medium 0.018), SBC KS p < 10−13. Split-conformal recalibration to 0.031; 90% intervals cover 0.76 → 0.88. Converged NICER flow: deviation 0.011, fails SBC on all 5 parameters.
- Gain marginalized
- g ~ U[0.95, 1.05]. Medium-count bias +0.018 ± 0.006 (fixed-gain flow) vs +0.020 ± 0.006 (marginalized), 0.03–0.08σ. Bright: +0.019, 0.05–0.11σ. Flow gain posterior 88–90% of the prior width. Exact gain-marginalized nested sampling on bright spectra shrinks the gain to 0.73–0.86 of the prior; the flow recovers 0.04–0.16 of that shrinkage.
Everything is simulation-based and no observed spectrum is fit. The gain error is injected as a slope only, with no offset. All results use one five-parameter source model, and the sub-percent gain sweep covers medium counts on EPIC-pn only.
Paper and code
Akbari, K. “Misspecification in amortized X-ray spectral inference: a detection benchmark, a gain shift three schemes cannot see, and marginalizing it out.” Preprint, arXiv:2606.17098.