Haroui Ma, Francesco Quinzan, Theresa Willem, Stefan Bauer
Abstract
Machine learning (ML) systems for medical imaging have demonstrated remarkable diagnostic capabilities. However, their susceptibility to learning spurious correlations with sensitive attributes poses significant risks to fairness and safety. In this paper, we introduce a novel statistical framework to evaluate the dependency of medical imaging ML models on sensitive attributes, such as demographics.
Introduction
A major problem in using machine learning (ML) systems for healthcare is that these systems exhibit biases leading to disparities in performance for different groups of patients [1–4]. For instance, patients from certain demographic groups may face a higher likelihood of misdiagnosis, such as false positives or false negatives. One possible reason for this problem is that AI models inadvertently learn and amplify existing confounding effects from the training data [5].
Methods
Overview
Our approach evaluates whether a medical AI model relies on sensitive attributes, such as the self-reported race or sex, when making predictions. The central idea is to ask a simple question: Would the model’s prediction change if the same patient had a different demographic attribute, while everything else remained identical? Since such data cannot be observed in reality, we use a generative framework to simulate it. First, each medical image is mapped into a latent representation.
Results.
Fig 2 reports the rejection rates of all methods across models grouped by their ECA values, with whiskers indicating variability across repetitions and significance level fixed at . Overall, CIT-LR exhibits the clearest alignment with the target notion of counterfactual invariance: for biased models ( ), it achieves rejection rates close to 100% in nearly all tasks, while DP and EO show significantly weaker and often non-monotone behavior, especially in the exponential, interactive, and log-exp settings.
Discussion
The proposed counterfactual invariance test via latent representations has significant implications for fairness in medical AI. The method explores the deep causal mechanism of the diagnostic models, revealing dependencies, which are overlooked by traditional association-based metrics. Through quantifying the biases linked to the sensitive attribute, our method identifies the diagnostic models’ undesirable dependencies across demographic groups, reducing the risk of misdiagnosis.
Conclusion
Acknowledgments
The authors gratefully acknowledge the Gauss Centre for Supercomputing e.V. (www.gauss-centre.eu) for providing computing time through the John von Neumann Institute for Computing (NIC) on the GCS Supercomputer JUPITER | JUWELS at Jülich Supercomputing Centre (JSC)
Citation: Ma H, Quinzan F, Willem T, Bauer S (2026) AI alignment in medical imaging: Unveiling hidden biases through counterfactual analysis. PLOS Digit Health 5(8): e0001516. https://doi.org/10.1371/journal.pdig.0001516
Editor: Onicio Batista Leal-Neto, The University of Arizona, UNITED STATES OF AMERICA
Received: July 2, 2025; Accepted: June 9, 2026; Published: August 13, 2026
Copyright: © 2026 Ma et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: The data used in this study includes both synthetic and publicly available real-world medical imaging datasets. Specifically, the CheXpert dataset is available from the Stanford Machine Learning Group at https://stanfordmlgroup.github.io/competitions/chexpert/ and the MIMIC-CXR dataset is available through PhysioNet at https://physionet.org/content/mimic-cxr/2.0.0/, subject to data use agreements. Synthetic data used for validation is available from the authors upon reasonable request. All code necessary to reproduce our experiments and analyses is available at https://github.com/Neferpitou3871/AI-Alignment-Medical-Imaging.
Funding: This work was partially supported by the Helmholtz Foundation Model Initiative, the Helmholtz Association, and by the German Federal Ministry of Education and Research (Grant: 01IS24082 to SB). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.
Competing interests: The authors have declared that no competing interests exist.