Skip to content

Repository files navigation

Galactic X-ray Source Classifier

A machine-learning pipeline that takes survey-level X-ray and Gaia features for eROSITA (eRASS1) sources toward the inner Galactic plane and returns calibrated class probabilities (star / CV / X-ray binary / AGN) for every source, including the ~20,750 in the footprint with no published classification. It covers all ~34k inner-plane sources, and a large part of the work is mapping where and why the classification fails.

It builds on the official eRASS1 counterpart catalogue (Salvato et al. 2025) and relates to the multiwavelength classification of Saeedi et al. 2024 and the inner-disc XMM classification of Bao et al. 2025. The probabilities are meant to sit alongside per-source spectral classification at survey scale, where source-by-source spectroscopy is not feasible.

The underlying question is how well broad Galactic-plane source classes separate using only survey-level X-ray and Gaia features with calibrated probabilities, and where that separation fails physically.


Cross-validation results

confusion matrix and reliability diagram

Left: row-normalized 4-class confusion matrix from leak-free 5-fold cross-validation (out-of-fold, calibrated argmax). Right: reliability diagram, raw vs temperature-calibrated, all four classes.

Per-class PR-AUC under 430:1 imbalance: STAR 1.00, AGN 0.96, XRB 0.76, CV 0.38. Two of these need a caveat. STAR's 1.00 sits only 0.028 above the no-skill base-rate floor of 0.972 (STAR is 97% of the labels), so the more informative STAR number is its ROC-AUC of 0.992. The AGN 0.96 is in-sample to the way AGN labels are defined: 86% of the AGN labels come from an exgal+p_any rule built on features the model also uses (p_any is the #1 permutation feature), so on the externally labeled (Milliquas-only) AGN the recall is 64%, against 97% on the in-sample route. See the failure modes section and Limitations.

The dominant off-diagonal flow runs into STAR (CV→STAR 40%, XRB→STAR 26%). Applied to the 20,750 unlabeled inner-plane sources, 93.2% receive a >90%-confident label, and the confident-AGN fraction rises from 0% in the mid-plane to ~3.3% at |b| = 5–10°, following the plane geometry. The confident assignments are star-dominated, which reflects where the model fails as much as the true population: in |b|<1° it assigns 96.7% STAR and 3.1% compact, while Bao et al. 2025 spectroscopy of the same inner-disc regime finds ~74% coronal and ~18% compact accretors, so the model recovers compact accretors about 6× too rarely. That star-dominated mix sets a lower bound on contamination of the coronal population. The detailed comparison is in the application section below and in Limitations.

The raw HistGradientBoosting probabilities are already close to calibrated (mean ECE ~1×10⁻³); temperature scaling and isotonic move it only at the fourth decimal, so calibration here confirms the probabilities are already close and does not move them materially. With n=30 CV, the rare-class reliability curves are dominated by sampling noise, so calibration is only testable in the bulk classes. Details in the Calibration section.


Prior work

This builds on a specific, recent literature. The contribution is a multi-class Galactic-population classifier (star / CV / XRB / AGN) with calibrated probabilities across the whole inner-plane footprint, finer-grained than the two-way Galactic/extragalactic split of the catalogues it builds on.

  • Salvato et al. 2025 (arXiv:2509.02842), the official NWAY-based Gaia DR3 counterpart catalogue for eRASS1, with a Galactic/extragalactic ML classification (the STAREX classifier). I build on its counterparts (using its p_any, the Gaia columns it carries, joined on its UID) and go finer-grained within the Galactic population instead of a two-way Galactic/extragalactic split. Its exgal flag is not held out as an independent cross-check: STAREX is trained on the same Gaia astrometry/photometry sibling features my model uses, and exgal+p_any is the rule that generated 86% of my AGN labels, so it functions as a label source here and cannot serve as an external validator (see the failure modes section). Their own held-out STAREX numbers (their Fig. 8) are EXGAL recall 99.2% / purity 93.3%, GAL recall 92.8% / purity 99.1%; my Gal/exgal-equivalent split is comparable but goes finer-grained (4 Galactic classes) and, unlike STAREX, is not spectroscopically validated.
  • Saeedi et al. 2024 (arXiv:2407.12583), multiwavelength classification of 8,311 faint eRASS1 sources in the Canis Major region. The closest methodological precedent; here I cover the full inner plane (|b|<10°, |l|≤60°) with calibrated probabilities, where Saeedi covers a single region. Saeedi computes f_X/f_opt with the (BP+RP)/2 + 5.37 optical convention; I use the Rodriguez 2025 Gaia-G convention. The two are offset conventions, so my diagnostic-plane separation lines are not directly comparable to Saeedi's.
  • Bao et al. 2025, "Unveiling the soft X-ray source population towards the inner Galactic disk with XMM-Newton" (arXiv:2510.23814), 189 sources, ~74% coronal / ~8% massive stars / ~18% compact accretors, with deeper XMM spectroscopy. This is the spectroscopic comparison sample for the inner-disc probabilities (see the application section). Bao's sample is the brighter, XMM-spectroscopable subset, so it is not the same flux-limited population I classify, but the direction and rough size of the compact-accretor under-recovery match the CV/XRB→STAR leak the classifier already shows in cross-validation.
  • Rodriguez et al. 2025 (arXiv:2408.16053), the "X-ray Main Sequence", the f_X/f_opt vs Gaia-color feature engineering I adopt exactly (§3.2.1).
  • Maccacaro et al. 1988 (1988ApJ...326..680M), the original f_X/f_opt vs optical-color diagnostic plane. My X-ray band is 0.2–2.3 keV (eROSITA ML_FLUX_1); Maccacaro used 0.3–3.5 keV. The diagnostic survives the band shift (Rodriguez 2025 §3.2.1).
  • Merloni et al. 2024 (arXiv:2401.17274), the eRASS1 DR1 catalogue itself.
  • Freund et al. 2024 (HamStar, J/A+A/684/A121), the coronal-star identifications that define my STAR labels (~91.5% reliability).
  • Avakyan et al. 2023 (LMXB, J/A+A/675/A199), Neumann et al. 2023 & Fortin et al. 2023 (HMXB), Fortin et al. 2024 (LMXB), the XRB label catalogues.
  • Ritter & Kolb final edition 7.24 (B/cb) + Schwope et al. 2024 (arXiv:2407.20903) eROSITA CV supplement, the CV labels. Schwope find only ~65% of known CVs are even detected in eRASS1, which sets the scale of the CV label scarcity below.
  • NWAY (Salvato & Buchner, arXiv:1705.10711), the Bayesian cross-match tool Salvato 2025 used; the counterpart-association method this repo relies on instead of reimplementing.
  • Milliquas v8 (Flesch, VII/294), the AGN labels (final edition; no newer version exists).

The diagnostic plane

fx_fopt vs BP-RP

The classical Maccacaro-style diagnostic, modernized with Gaia: log₁₀(F_X/F_opt) vs Gaia BP−RP, colored by class, with the Rodriguez 2025 empirical separation lines overlaid. Coronal stars sit low (optically bright relative to their X-rays); CVs and XRBs lift up and to the right; AGN scatter across it because most have no useful Gaia counterpart. Permutation importance later confirms the model leans on these axes (p_any, HR3, log_fx_fopt).


What's in here

galactic-xray-source-classifier/
  README.md
  requirements.txt
  LICENSE
  configs/        dataset.yaml, features.yaml, model.yaml  (every run seeded)
  src/gxclass/    datasets.py, labels.py, features.py, model.py
  scripts/        download_data.py, build_dataset.py, build_features.py,
                  train_model.py, classify_unlabeled.py, make_plots.py, run_all.py
  notebooks/      walkthrough.ipynb   (end-to-end, executed)
  tests/          pytest; HR formulas, calibration math, OOD flag, accounting
  outputs/        committed: money_plot.png, classified_sample.{parquet,csv},
                  diagnostics/  (confusion, reliability, permutation, fx_fopt)
  data/           NOT committed; rebuilt by scripts/download_data.py
  docs/column_map.md   verified FITS column names for every catalogue

Everything in outputs/ is committed so the numbers can be read without running the pipeline. data/ is not committed and is rebuilt by the download/build scripts.


Results

The footprint is |b| < 10°, |l| ≤ 60°. The Phase-1 funnel: 930,203 eRASS1 Main sources → 903,521 point sources → 34,093 in the inner-plane footprint, of which 30,216 have a Salvato Gaia counterpart. Label crossmatch gives 12,965 STAR / 30 CV / 53 LMXB / 25 HMXB / 270 AGN, leaving 20,750 unlabeled.

The 4-class decision

With only 25 HMXB in the footprint (below the ~30 gate I set in advance), I merge LMXB+HMXB into a single XRB class for the 4-class model (STAR / CV / XRB / AGN) and keep the 5-class split as a documented appendix. The threshold was fixed from the label counts beforehand. The 5-class run confirms the LMXB↔HMXB boundary is not reliably learnable at 25+53 labels (HMXB→LMXB leak 24%; HMXB recall 44%).

Per-class metrics (out-of-fold, calibrated)

class support precision recall F1 ROC-AUC PR-AUC
STAR 12,965 1.00 1.00 1.00 0.992 1.00 (base-rate floor 0.972)
CV 30 0.56 0.30 0.39 0.96 0.38
XRB 78 0.81 0.67 0.73 0.98 0.76
AGN (in-sample to the exgal+p_any label rule, circular; see Limitations) 231 0.93 0.97 0.95 1.00 0.96
AGN (Milliquas-only, external labels) 39 n/a 0.64 n/a 0.99 0.57
AGN (all labels combined) 270 0.93 0.92 0.92 1.00 0.96

PR-AUC is the appropriate metric under the 430:1 STAR:CV imbalance, except for STAR, whose PR-AUC of 1.00 sits only 0.028 above the no-skill base-rate floor of 0.972 (STAR is 97% of the labels); for STAR the ROC-AUC of 0.992 is the informative number.

The AGN rows are split by label provenance because the two routes are not the same evidence. 86% of the AGN labels (231/270) were defined by the exgal+p_any≥0.5 rule, and p_any is a model feature (the #1 permutation feature), so the AGN recall of 97% on that route is in-sample to the label rule: the model is partly scored on a label it helped define. On the externally labeled AGN (Milliquas-only, 39 sources, an independent catalogue that does not use p_any), recall drops to 64% (25/39) and the OvR PR-AUC to 0.57. That 33-point recall gap measures the circularity directly (see the failure modes section and Limitations).

De-circularized ("no-counterpart") retrain

To quantify the leak, I retrain a variant (configs/model_nocp.yaml) that excludes the counterpart-derived features (p_any, p_i, nway_separation, has_counterpart + their missing-indicators; 39 → 31 features) under an otherwise identical CV/calibration protocol. The no-counterpart numbers stand alongside the main model:

class PR-AUC (main → nocp) recall (main → nocp)
STAR 1.00 → 1.00 1.00 → 1.00
CV 0.38 → 0.37 0.30 → 0.30
XRB 0.76 → 0.73 0.67 → 0.68
AGN (all) 0.96 → 0.92 0.92 → 0.87
AGN exgal-route (circular) n/a 0.97 → 0.91
AGN Milliquas-only (external) 0.57 → 0.48 0.64 → 0.64

Removing the label-generating feature costs 4.8 recall points on AGN, borne almost entirely by the exgal-route labels (0.97 → 0.91); STAR/CV/XRB barely move because their discriminators are the hardness ratios and log_fx_fopt. The external Milliquas recall is identical (0.64) in both models, which was the circularity-free number throughout; its ranking (PR-AUC 0.57 → 0.48) shows p_any was doing partly circular ranking work. The de-circularized AGN result is recall 0.87 combined and 0.64 external. Full table and interpretation in RESULTS.md.

AGN and XRB are otherwise well-recovered. CV is the hard class (PR-AUC 0.38, recall 0.30): only 30 labels, and CVs sit in the same hardness / f_X/f_opt region as the overwhelming coronal-star sea. The primary model is HistGradientBoostingClassifier (native NaN handling, so missing values are kept as a signal instead of imputed; balanced sample weights); a median-impute RandomForest baseline is slightly worse across the board (CV PR-AUC 0.23 vs 0.38), consistent with the native-NaN advantage.

Calibration

Per-class one-vs-rest, leak-free: out-of-fold raw probabilities are calibrated by a second independent stratified 5-fold CV, so no row is calibrated by a calibrator that saw it. Two calibrators compared (isotonic and temperature scaling) on adaptive-binned ECE:

mean ECE (over 4 classes)
raw (uncalibrated) 0.00112
isotonic 0.00144
temperature (winner) 0.00101

Temperature scaling wins, but the absolute ECE numbers are tiny (~10⁻³) and the calibrators differ only at the fourth decimal. This follows from the imbalance: each per-class OvR confidence is dominated by the ~13k true-negative rows clustered near the class base rate, where the model is already well-calibrated. Calibration measurably helps on a controlled synthetic overconfident toy (a test asserts both methods drive ECE from >0.03 toward ~0), but on this dataset the raw HGB was already close to calibrated.

Permutation importance

Out-of-fold, 15 repeats, CV error bars. Top 5: p_any (0.14), HR3 (0.08), log_fx_fopt (0.08), HR3_err (0.07), HR2 (0.07). The counterpart-match probability dominates. Two qualifications on the "model rediscovered the textbook discriminators" reading.

p_any is not a physical axis. It is NWAY's counterpart match-probability, a measure of whether a Gaia source sits at the X-ray position. It tops the list largely because it is also half of the rule that generated the AGN labels (exgal+p_any). Its dominance comes partly from the labelling itself, and only partly from source astrophysics.

The physical discriminators are the ones below it: HR3/HR2 (X-ray hardness) and log_fx_fopt (the classical Maccacaro 1988 X-ray/optical flux ratio). The model leans on these, and the CV/XRB↔STAR confusion plays out along them.

The same caution applies to STAR: the HamStar coronal labels are built from Gaia counterpart/astrometry properties the model also ingests, so the near-perfect STAR ROC-AUC partly reflects the model recovering its own label-construction features (a weaker version of the AGN circularity). The *_missing indicator columns carry near-zero permutation importance for this label set.

Applying it to the 20,750 unlabeled sources

Refit the final 4-class HGB on all labeled sources, calibrate with the same leak-free temperature scaling fit on the out-of-fold probabilities (the calibrator never sees an unlabeled row), and predict:

  • 93.2% (19,331 / 20,750) of the unlabeled tail receives a >90%-confident label. Here "confident" means confidence under extrapolation: calibration is validated only in the bright labeled regime (median G 13.7), while the unlabeled tail is ~6.4 mag fainter (median G 20.1) with far weaker counterparts (median p_any 0.62 → 0.015). See the covariate-shift table in Limitations. The OOD flag does not rescue this: the confident-label rate is essentially the same inside the OOD flag (89.5%) as outside it (93.6%), because the shift is a global property of the unlabeled set; the flag is built to catch per-row outliers instead.
  • The confident population is star-dominated: 97.7% STAR, 1.9% AGN, 0.3% XRB, 0.1% CV. This reflects the direction the classifier fails in. In |b|<1° the model assigns 96.7% STAR / 3.1% compact; Bao et al. 2025 spectroscopy of the same inner-disc regime finds ~74% coronal / ~18% compact accretors, so the model recovers compact accretors ~6× too rarely (with the brighter-sample caveat noted in Prior work). The star-dominated mix sets a lower bound on the contamination of the coronal-star population by faint accretors.
  • AGN fraction rises with |b|, monotonically: 0.00% at |b|<1°, 0.00% at 1–2°, 0.35% at 2–5°, 3.27% at 5–10°. Cleaner extragalactic sight lines off the crowded, absorbed mid-plane. A basic sanity check the model passes (the AGN identification itself inherits the in-sample circularity discussed in the failure modes section).
  • 2,262 sources are OOD-flagged (a core feature outside the [0.5, 99.5] training percentile band, or an unseen missing-value pattern). Their probabilities should be read with extra caution. This flag catches per-row outliers, not the global bright→faint covariate shift.

Sanity checks

  • Star-dominated mix on plane sight lines: yes. 97.7% of confident assignments are STAR, as expected looking through the disc.
  • AGN fraction vs latitude: rises monotonically with |b| (above), as expected.
  • HamStar consistency: every STAR label is a HamStar coronal identification by construction, so there are zero HamStar-coronal sources left in the unlabeled set; the naive confident-STAR-vs-HamStar overlap is therefore empty, which is expected. The cross-check against Salvato's exgal flag (0.56% of confident-STAR are exgal=True, close to zero as expected) is only a consistency check: exgal is not a model feature, but it is built from the same Gaia sibling features the model uses and supplies 86% of the AGN labels, so it cannot validate the model from outside. The confident-AGN vs exgal disagreement is discussed in the failure modes section.

The flags are simple and explicit: low_confidence = top calibrated prob < 0.9; ood_flag = feature-space novelty as above. No new heavy dependencies.

The full table is outputs/classified_sample.parquet (and a full CSV): source ID, position (RA/Dec, l/b), the four calibrated class probabilities, top class, top prob, and the low_confidence / ood_flag columns.


Failure modes

What matters here is which classes confuse, and why physically. Using the actual numbers:

AGN label–feature circularity. This concerns what the AGN number means. 86% of the AGN labels (231/270) come from the exgal+p_any≥0.5 rule, and p_any (with its siblings p_i, nway_separation, has_counterpart) are model features (p_any is the #1 permutation feature). So when the model scores AGN recall 97% on the exgal-route AGN, it is partly graded on a label it helped define: p_any → AGN label → p_any feature is a closed loop. On the external subset (39 Milliquas-only AGN, an independent catalogue that does not use p_any), recall is 64% (25/39) and OvR PR-AUC 0.57, 33 recall points and 0.39 PR-AUC below the in-sample route. The in-sample AGN 0.96 is an upper bound set by the label rule; the externally validated numbers are the 0.57/0.64 ones. A weaker version affects STAR: the HamStar coronal labels are built from Gaia counterpart/astrometry properties the model also uses, so the near-perfect STAR ROC-AUC partly reflects the model recovering its own label-construction features.

CV→STAR (40%, 12/30 leak; recall 30%) and XRB→STAR (26%): faint hard sources leak into the coronal sea. This is the dominant off-diagonal flow, and it is physical. Quiescent X-ray binaries and non-magnetic CVs overlap coronally-active stars in exactly the features I have: hardness ratio and f_X/f_opt. A faint quiescent XRB with a bright optical counterpart looks, to survey-level features, like an active star. At 30 CV and 78 XRB labels against ~13k stars, the model's safest bet for a borderline source is STAR, and it takes it. Spectroscopy (orbital modulation, emission lines, X-ray spectral shape) breaks this degeneracy; survey photometry alone does not.

LMXB↔HMXB (the 5-class appendix): HMXB→LMXB 24%, LMXB→HMXB 9%. With 25 HMXB and 53 LMXB sharing overlapping survey X-ray + Gaia features, the donor-mass distinction (which is what L/H-MXB means) is not encoded strongly enough in eROSITA hardness + Gaia color to separate at this label volume. HMXB is the bigger loser (recall 44%). This is why I merged them for the 4-class model. The 4-class XRB (F1 0.73) recovers what the split loses (LMXB F1 0.58, HMXB F1 0.46).

Magnetic-CV / quiescent-XRB hardness + f_X/f_opt overlap. The same region of the diagnostic plane that catches CVs also catches the hardest stars and the softest XRBs. The diagnostic plane figure above shows the overlap directly: there is no clean line, only a gradient.

AGN the model finds vs AGN Salvato's classifier keeps Galactic. Of the 367 confident-AGN unlabeled sources, only 4.6% are flagged exgal=True by Salvato's STAREX Galactic/extragalactic ML classifier. These are not obvious stars: median parallax S/N ≈ −0.08 and proper-motion S/N ≈ 0.84 (astrometrically unconstrained), faint (G≈19.4), red, with AGN-like f_X/f_opt. This is a disagreement between two classifiers on faint plane sources, and the two are not independent: STAREX shares Gaia sibling features with my model, so they are correlated estimators. My model calls them AGN from X-ray hardness + flux ratio + low counterpart confidence; STAREX, astrometry-weighted with Galactic priors at low |b|, keeps them Galactic. Resolving them needs spectroscopy, which is the representativeness tension that motivates the Bao 2025 XMM follow-up.

Label-noise candidates. The top-20 most confidently-rejected training labels (outputs/diagnostics/label_noise_top20.csv) are overwhelmingly faint hard XRB/CV labels from heterogeneous catalogues whose eROSITA counterpart looks coronal. Several have p_label ≈ 0 and p_top_other ≈ 1.0 (e.g. 1eRASS J175141.7−351643, labeled CV, predicted STAR at p=0.9985). Ritter-Kolb CV coordinates are heterogeneous literature positions, so these are the prime candidates for a wrong-X-ray-counterpart check. This is a candidate list for follow-up; some are genuinely hard sources, not all are mislabels.


Limitations

  1. AGN label–feature circularity (leakage), quantified by a retrain, with a weaker version for STAR. 86% of AGN labels are defined by the exgal+p_any≥0.5 rule, and p_any is the #1 model feature, so the loop is p_any → AGN label → p_any feature. The in-sample AGN metrics (recall 97%, PR-AUC 0.96) overstate what the model has learned; the externally validated number, on the 39 Milliquas-only AGN, is recall 64% / PR-AUC 0.57. The de-circularized variant (configs/model_nocp.yaml), which drops the counterpart-derived features, lowers combined AGN recall from 0.92 to 0.87 (PR-AUC 0.96 to 0.92), concentrated on the exgal-route labels (0.97 to 0.91), while STAR/CV/XRB barely move and the external Milliquas recall stays at 0.64 (see the de-circularized table above and RESULTS.md). STAR has the same issue one notch weaker: HamStar coronal labels are built from Gaia counterpart/astrometry features the model also ingests, so the near-perfect STAR ROC-AUC partly measures the model recovering its own label-construction inputs. Metrics that are in-sample to their label rule are flagged as such in the tables above.

  2. Covariate shift: the labels are bright, the targets are faint. The calibration is fit and validated on the labeled set, which is systematically brighter and better-counterparted than the 20,750 unlabeled sources it is then applied to. The 93.2% confident fraction is confidence under extrapolation:

    feature labeled (median) unlabeled (median)
    Gaia G 13.67 20.10
    p_any (counterpart match prob) 0.62 0.015
    DET_LIKE 13.2 7.6
    log f_X/f_opt −3.13 −0.79
    HR2 −0.23 +0.09

    The marginal-band OOD flag does not capture this: the confident-label rate is 89.5% inside the flag vs 93.6% outside it (about the same), because the shift is a global property of the unlabeled set; the flag is built to catch per-row outliers instead. I keep the flag (it does catch genuine per-feature outliers and unseen missingness patterns), but it is blind to the dominant bright→faint shift. The unlabeled probabilities are extrapolations whose calibration is verified only in the bright regime.

  3. Tiny CV and XRB label sets. 30 CVs and 78 XRBs in-footprint. The CV PR-AUC of 0.38 is what 30 labels buys you. Schwope 2024 found only ~65% of known CVs are even detected in eRASS1, so this scarcity is structural and not fixable by better crossmatching. With n=30, the calibration of the rare-class probabilities is effectively untestable: the rare-class reliability curve is sampling noise.

  4. Label provenance is heterogeneous. STAR (HamStar coronal), CV (Ritter-Kolb + Schwope), XRB (Avakyan/Neumann/Fortin), AGN (Milliquas + Salvato exgal) come from catalogues with different selection functions, coordinate precision, and reliability. The label-noise pass surfaces the worst cases but cannot remove the underlying heterogeneity.

  5. eRASS1-depth censoring. The plane is absorbed and crowded; soft (P1) and hard (P4) bands are often undetected (HR1 missing 52%, HR3 55%). The model leans on HR2 and p_any, which survive, but it is fundamentally a shallow-survey classifier.

  6. STAR labels = HamStar coronal, by construction. Every STAR label is a HamStar coronal identification, so a naive "compare confident-STAR against HamStar" overlap check is empty on the unlabeled set. The exgal cross-check shares features with the model and is itself a label source, so it serves as a consistency check only (see item 1).

  7. Calibration verified, with small effect. The calibration machinery is correct and leak-free, but on this imbalanced dataset it moves ECE at the fourth decimal. The bulk-class probabilities check out; in the rare classes it is untestable.

  8. Not a spectral classifier. Every probability comes from survey photometry and Gaia astrometry. It operates at survey scale, where per-source spectroscopy is not available; the Bao 2025 comparison (~6× under-recovery of compact accretors in |b|<1°) is the quantitative statement of this limit.


Quickstart and reproducibility

python -m venv .venv && .venv/Scripts/python -m pip install -r requirements.txt

# rebuild data/ from public archives (resumable, cached) and run everything:
python scripts/run_all.py --from-scratch

# or, with the processed parquets already present, just the modelling half:
python scripts/run_all.py

# individual stages (each standalone):
python scripts/download_data.py    --only all
python scripts/build_dataset.py    --config configs/dataset.yaml
python scripts/build_features.py   --config configs/features.yaml
python scripts/train_model.py      --config configs/model.yaml
python scripts/classify_unlabeled.py --config configs/model.yaml
python scripts/make_plots.py       --config configs/model.yaml

Every run is seeded (seed: 42 in the configs). The committed figures and classified_sample can be read directly from outputs/. Regenerating them requires the processed parquets, which are produced by the pipeline: download_data.py fetches the public inputs (eRASS1 Main v1.2, the Salvato GDR3 counterpart catalogue, the VizieR label catalogues, HamStar, and the public Milliquas v8 catalogue, Flesch VizieR VII/294), build_dataset.py and build_features.py assemble data/processed/, and train_model.py / classify_unlabeled.py / make_plots.py produce the model outputs and figures. data/ is not committed. Config files carry inline citation comments for every adopted constant. sklearn parallelism is capped at n_jobs: 6.

pytest covers the physics-critical and ML-critical functions: hardness-ratio formulas and error propagation, f_X/f_opt against hand-computed values, ECE and calibration math, the leak-free CV split, the OOD flag, and the confident-fraction accounting. 90 tests pass.


Author

Karan Akbari, MSc Astrophysics, St. Xavier's College Mumbai. Background in X-ray timing and spectral analysis of black-hole X-ray binaries (GRS 1915+105, 4U 1630−47) with Dr. Sudip Bhattacharyya at TIFR.

About

Calibrated gradient-boosted classifier (star, CV, X-ray binary, AGN) for eROSITA Galactic-plane sources with Gaia counterparts, with leakage-free calibration and out-of-distribution flagging.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages