datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wds_carsnpm3d-kitti-carlapdm_carlaCyclePrefDB-I2T-Reconstructions
Image Reconstructions for CyclePrefDB-I2T
Project page | Paper | Code
This dataset contains reconstruction images used to determine cycle consistency preferences for CyclePrefDB-I2T. You can find the corresponding file paths in the CyclePrefDB-I2T dataset here. Reconstructions are created using Stable Diffusion 3 Medium.
Preparing the reconstructions
You can download the test and validation split .tar files and extract them directly.Use this script to extract the… See the full description on the dataset page: https://huggingface.co/datasets/carolineec/CyclePrefDB-I2T-Reconstructions.cardcaptor-merged-dataset
CardCaptor Trading Card Dataset · merged_dataset_50x3
This is the dataset used to train the CardCaptor OBB Detector v3 model.
📦 Contents
~60,408 synthetic training renderings (dataset_0)
~220 real captures × 50× deterministic photometric multiplier ⇒ 11,000 effective handcrafted training images
Validation: 55 handcrafted frames (210 ground-truth oriented boxes) to mirror real-world usage
Format: YOLO OBB (class cx cy w h x1 y1 x2 y2 x3 y3 x4 y4) under images/ + labels/… See the full description on the dataset page: https://huggingface.co/datasets/AlecKarfonta/cardcaptor-merged-dataset.Stanford-Carsscryfall-card-imagescarol-subcorpora
Subcorpora Carolina
Contains Carol·B and Carol·(D+B), respectively balanced and deduplicated subcorpora of the Carolina Corpus (Bea) version.
Carol·B was balanced in terms of tokens per domain from Carolina Corpus. This means that it contains approximately the same number of tokens (~60,2M) from each of Carolina’s largest domains: Legislative, Instructional, Entertainment, Journalistic, Juridical and Virtual Forum. Carol·B has, in total, 361,071,147 tokens and 5,5 GB.
Carol·(D+B)… See the full description on the dataset page: https://huggingface.co/datasets/carolina-c4ai/carol-subcorpora.wds_carscardano_cipsphysense_carla_dataset
PhySense CARLA Synthesized Dataset
Dataset Summary
This dataset is the CARLA-synthesized traffic dataset used in PhySense: Defending Physically Realizable Attacks for Autonomous Systems via Consistency Reasoning, CCS ’24. It is generated using the CARLA simulator and CARLA PythonAPI.
Source / Collection
The dataset is synthesized in CARLA. Instructions and scripts to reproduce/collect the dataset are provided in the accompanying GitHub repository:
Collection… See the full description on the dataset page: https://huggingface.co/datasets/Ruoyao/physense_carla_dataset.ML2025_HW_Cardiac_Muscle
ML 2025 Cardiac Muscle Dataset
You have to extract the file before using the dataset.
carlaoo3d_toy_car
