astro-ph.IMJun 17, 2026

The Chandra-Gaia Catalog of Counterparts: Resolving ambiguous Gaia matches to X-ray sources in the Chandra Source Catalog using Machine Learning

Authors: V. Samuel Pérez-DíazVinay L. KashyapJoshua D. IngramDavid FouheyJuan Rafael Martínez-GalarzaPavlos ProtopapasJeremy J. DrakeDong-Woo Kim+1 more

Organizations: Center for Astrophysics | Harvard & Smithsonian, 60 Garden St, Cambridge MA 02138, USA · Harvard John A. Paulson School of Engineering and Applied Sciences, 150 Western Ave, Allston, MA 02134, USA · Universidad del Rosario, School of Engineering, Science and Technology, Cll. 12C No. 6-25, Bogot´a, Colombia · The NSF AI Institute for Artificial Intelligence and Fundamental Interactions, Cambridge MA 02139, USA · New York University, Courant Institute, 60 5th Avenue, New York NY, USA · Carnegie Mellon University, 5000 Forbes Avenue, Pittsburgh, PA 15213 · New College of Florida, 5800 Bayshore Road, Sarasota, FL 34243, USA · Lockheed Martin Solar and Astrophysics Laboratory, 3251 Hanover St, Palo Alto, CA 94304, USA

Abstract

We present a framework to cross-match sources from the Chandra Source Catalog (CSC v2.1) with optical sources from Gaia Data Release 3. Unlike purely spatial approaches, we use source properties such as magnitudes, colors, and distances to identify true counterparts, detect chance coincidences, and resolve ambiguities when multiple plausible candidates exist. We define a training set of high-confidence matches using NWAY, a Bayesian cross-matching framework that accounts for positional errors and source densities. We train a gradient-boosted classifier (LightGBM) on a variety of features from both catalogs. Of the ~254254k unique X-ray sources, we find counterparts for ~113113k sources, of which plausible multiple counterparts are found for ~77k. We find no counterparts for ~2020k sources for which separation-based cross-matching does find a match, and attribute half of these to chance coincidences. We validate the pipeline on the Chandra Orion Ultradeep Project (COUP), where the machine-learning matches reproduce 95% of NWAY cross-matches without using any positional information. We release a catalog of the ~113113k Chandra-Gaia counterparts, together with ~77k alternative matches and ~2020k ambiguous NWAY associations, supporting future population studies of sources detectable by both Chandra and Gaia. We discuss limitations and provide a generalization of the framework that is applicable in other cross-matching scenarios.

Explore similar work

CardsList