qureaiorg/ct2xr-projections
Component-wise translated CT projections Synthetic chest radiographs derived from 21,887 chest CT volumes of CT-RATE, released as the translated outputs of the pipeline described in Anatomy-Decomposed Chest Computed Tomography (CT) Projections as Scalable Supervision for Bone Suppression in Chest Radiographs Angaitkar, Kumar, Satia, Rao, Mittal, Tadepalli, Putha — arXiv:2609.24937 (2026; under review at Medical Image Analysis). Released by Qure.ai. Version 1.0.0 (2026-09-15).… See the full description on the dataset page: https://huggingface.co/datasets/qureaiorg/ct2xr-projections.
Component-wise translated CT projections
Synthetic chest radiographs derived from 21,887 chest CT volumes of CT-RATE, released as the translated outputs of the pipeline described in
Anatomy-Decomposed Chest Computed Tomography (CT) Projections as Scalable Supervision for Bone Suppression in Chest Radiographs Angaitkar, Kumar, Satia, Rao, Mittal, Tadepalli, Putha — arXiv:2609.24937 (2026; under review at Medical Image Analysis).
Released by Qure.ai. Version 1.0.0 (2026-09-15). For reproducible use, pin a repository revision (revision= in load_dataset). Questions: the Community tab of this repository, or mrunmay.angaitkar@qure.ai.
Each CT was decomposed into anatomical components, each component was projected to a digitally reconstructed radiograph (DRR), and the bone and soft-tissue projections were then passed separately through an unpaired image-to-image translation designed to preserve anatomical content while adapting the appearance toward that of real chest radiographs. The translated components were recombined with the untranslated lung projection into the final synthetic radiograph. This release contains the two translated components, the lung projection, and the recombined radiograph.
What a row contains
One row = one source CT, projected PA at 100 kVp, four pixel-aligned 512×512 8-bit grayscale images:
The three component images are the aligned inputs used to construct the supplied translated radiograph; recombination applies weighting and normalisation, so the components are attributable, aligned constituents of that image rather than a plain pixel-wise sum. translated is the image evaluated in the manuscript's realism and anatomical-fidelity comparison (manuscript-reported FID 8.2 vs VinDr-CXR and 16.2 vs CheXpert, computed in a chest-radiograph-pretrained feature space, not the standard Inception-feature FID).
Labels
Six binary findings (int8, 0/1) copied from the source CT's CT-RATE multi-abnormality labels (every released volume has an explicit 0 or 1 for all six in the source metadata; nothing was imputed), which are derived from the CT report: lung_opacity, pleural_effusion, atelectasis, cardiomegaly, consolidation, lung_lesion. A 0 means the finding was not reported for the CT, not that its absence was radiographically confirmed, and a finding reported in the CT may not be visible in a PA projection. In the manuscript, four of these have definitions that match between CT and chest-radiograph label sets (pleural_effusion, atelectasis, cardiomegaly, consolidation); lung_opacity and lung_lesion denote broader concepts in the CT labels (the former includes ground-glass change, the latter fissural nodules) and should not be read as chest-radiograph ground truth.
Source-CT metadata
spacing_x (mm, in-plane voxel spacing), spacing_z (mm, slice spacing), rows, columns, num_slices describe the source CT volume. They are not the physical spacing of the released PNGs, which are in the projector's square 512×512 frame. kvp is the simulated tube potential (100).
Split and grouping
The single train split is a packaging split containing the whole release. It is neither the training cohort of the manuscript's suppression models nor a prescribed evaluation partition; make your own partitions. volume_name follows CT-RATE's {origin}_{patient}_{scan}_{reconstruction} pattern (e.g. train_1_a_2), where origin is the CT-RATE split the volume came from (20,609 from train, 1,278 from valid). The release covers 14,374 patients; 5,181 of them contribute more than one volume (up to 25). Group by the first two tokens (train_1) to avoid patient overlap across your partitions.
Usage
Tested with datasets 5.0.1, pyarrow 23.0.1, huggingface_hub 0.31.4, Pillow 11.3.0 (pip install datasets pillow).
from datasets import load_dataset
ds = load_dataset("qureaiorg/ct2xr-projections", split="train") # ~7.8 GB
r = ds[0]
r["bone_translated"], r["soft_translated"], r["lung"], r["translated"] # PIL images, 512x512, aligned
r["volume_name"], r["cardiomegaly"]
# patient id for grouping
patient = "_".join(r["volume_name"].split("_")[:2])Stream it instead of downloading:
ds = load_dataset("qureaiorg/ct2xr-projections", split="train", streaming=True)
r = next(iter(ds))Uses
Realistic synthetic chest radiographs with a known, aligned component decomposition: each translated radiograph comes with the bone and soft-tissue components it was assembled from. Intended uses are realism and anatomy evaluation of synthetic radiographs, domain-adaptation research, and qualitative comparison against other DRR renderers. Use as training augmentation is a proposed research direction; the manuscript does not establish a downstream augmentation benefit.
How it was made
- Segment each CT into bone, lung and other soft tissue.
- Project each component separately with a polychromatic attenuation model through one fixed PA geometry, so all components of a CT share a pixel grid exactly.
- Translate the bone and soft-tissue projections separately with an unpaired image-to-image model toward the appearance of the corresponding structures in real radiographs, then recombine them with the lung projection.
Source CTs are the CT-RATE volumes with axial slice spacing ≤ 1 mm. Images are in the projector's square frame (the body fills the frame at ~1:1), not a physical detector frame.
Limitations
- Synthetic. These are translated projections, not acquired radiographs. They carry CT-derived anatomy, the projector's assumptions and the translator's texture; the manuscript reports population-level label readability and lung-field geometry, not preservation of every individual finding.
- Translated components only. The untranslated bone and soft-tissue projections are not included; the lung projection is, because it is part of the recombination.
- Adults, frontal, one geometry. Only the 100 kVp PA projection of each CT is included.
- Labels are report-derived (see Labels above).
- Research use only. Not a medical device.
License
Released under CC BY-NC-SA 4.0. Every image here is derived from CT-RATE (CC BY-NC-SA 4.0), whose terms permit sharing of derived data under the same licence and prohibit commercial use and redistribution of the source data; this release contains no CT volumes and no reports. The label and metadata columns reproduce per-volume fields of CT-RATE's public metadata for the volumes used. Non-commercial, share-alike, attribution required. For commercial use of CT-RATE-derived data, contact the CT-RATE authors.
Changelog
- 1.0.0 (2026-09-16): initial public release; Parquet shards written with small row groups and a page index so the Hub viewer can page through rows.
Citation
If you use this dataset, please cite the paper (journal reference will be added on acceptance):
@misc{angaitkar2026anatomydecomposedchestcomputedtomography,
title={Anatomy-Decomposed Chest Computed Tomography (CT) Projections as Scalable Supervision for Bone Suppression in Chest Radiographs},
author={Mrunmay Angaitkar and Piyush Kumar and Aarjav Satia and Pranav Rao and Ashish Mittal and Manoj Tadepalli and Preetham Putha},
year={2026},
eprint={2609.24937},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.24937},
}Please also cite CT-RATE, from which every image here is derived:
@article{hamamci2026generalist,
title = {Generalist foundation models from a multimodal dataset for 3D computed tomography},
author = {Hamamci, Ibrahim Ethem and Er, Sezgin and Wang, Chenyu and others},
journal = {Nature Biomedical Engineering},
year = {2026}
}