datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CarlaOcc
Database_structure
CarlaOcc/
├── CarlaOccV1/
│ ├── calib/
│ │ └── calib.yaml
│ ├── splits/
│ │ ├── test.txt
│ │ ├── train.txt
│ │ └── val.txt
│ ├── SceneMeshes/
│ │ ├── fg_actors/
│ │ ├── fg_actor_occ/
│ │ └── TownXX_Opt/
│ │ ├── bg_actors/
│ │ └── bg_actor_occ/
│ ├── TownXX_Opt_SeqXX/
│ │ ├── poses/
│ │ │ ├── cam_00.txt
│ │ │ └── lidar.txt
│ │ ├── rgb/
│ │ │ ├── image_00/
│ │ │ │ ├── 0000.png… See the full description on the dataset page: https://huggingface.co/datasets/fengyi233/CarlaOcc.anemia-survey-dataset
Anemia Detection — Multi-Modal Clinical SEWA Rural Dataset
Organisation: SEWA Rural — Society for Education, Welfare and Action (Rural), Jhagadia, Gujarat, India
Dataset: sewa-rural-care/anemia-survey-dataset
Contact: sewarural@ymail.com
Version: 1.0 — July 2026
Dataset Summary
This dataset supports research into non-invasive, smartphone-based anemia
screening applicable to low-resource and rural healthcare settings. It was
collected by SEWA Rural — a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/sewa-rural-care/anemia-survey-dataset.carla-autopilot-multimodal-dataset
CARLA Autopilot Multimodal Dataset
This dataset contains synchronized multimodal driving data collected in the CARLA simulator using the autopilot feature. It provides RGB images from multiple cameras, semantic segmentation, LiDAR point clouds, 2D bounding boxes, and ego-vehicle state/control signals across varied weather, maps, and traffic densities.
The dataset is designed for research in autonomous driving, sensor fusion, imitation learning, and self-driving evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/immanuelpeter/carla-autopilot-multimodal-dataset.cardiac_cine_acdc
ACDC (Cardiac Cine-MRI)
ACDC (Automatic Cardiac Diagnosis Challenge, MICCAI 2017) is a cine‑MRI dataset for cardiac segmentation.This repository contains processed NIfTI files in Data/processed_output/acdc format.
Dataset Summary
Modality: Cardiac cine‑MRI (NIfTI)
Task: Segmentation of LV, RV, and myocardium
Frames: ED/ES + full SAX time series (sax_t)
Labels: LV/RV cavities + myocardium
Splits: train, test (as provided in processed output)
Data Structure (per… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_acdc.toy-car-annotation-YOLOHey everyone,
In my final year project, I created Smart Traffic Management System.The project was to manage traffic lights' delays based on the number of vehicles on road.I made everything worked using Raspberry Pi and pre-recorded videos but it was a "final year project", it was needed to be tested by changing videos frequently which was a kind of hustle. Collecting tons of videos and loading them in Pi was not too hard but it would have cost time, by every time changing names of videos in… See the full description on the dataset page: https://huggingface.co/datasets/tubasid/toy-car-annotation-YOLO.lalafo-kg-cars
lalafo.kg — Kyrgyzstan Cars (used-car market)
Scraped from lalafo.kg, the largest informal classifieds
board in Kyrgyzstan — messier and larger than the curated boards, and closer to the
real street-level market. Field names are English; values are kept in the original
language (Russian).
Subsets
subset
rows
description
listings
54,518
one row per advertisement (every category, deal and region in scope)
users
49,685
sellers (ad authors), with… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/lalafo-kg-cars.pokemon_card_image_for_authenticity_classification
Pokemon Card Image for Authenticity Classification
This dataset contains front/back images of Pokemon cards for authenticity experiments.
Dataset structure
Images/: all image files (.jpeg)
Images/metadata.jsonl: metadata used by Hugging Face imagefolder
labels.csv: flat label file with the same rows as metadata
Columns
image: image object loaded from file
id: image filename (unique id)
side: card side (0 = front, 1 = back)
labels: authenticity label (1 =… See the full description on the dataset page: https://huggingface.co/datasets/stevelohwc/pokemon_card_image_for_authenticity_classification.index-cards-cas-galapagos-stewart-specimens
Alban Stewart's Galápagos Expedition Specimen Cards (CAS Archives, 1905–1906)
1,059 specimen index cards from the California Academy of Sciences Archives,
documenting Alban Stewart's botanical specimens collected on the 1905–1906
California Academy of Sciences Galápagos Expedition. One card per specimen with
species, locality on the islands, collection date and field notes; companion to
Stewart's expedition journal.
Harvested from the 6 IA items
csfa788562626 +
csfa788562626Alpha +… See the full description on the dataset page: https://huggingface.co/datasets/biglam/index-cards-cas-galapagos-stewart-specimens.stanford-cars-lance
Stanford Cars (Lance Format)
A Lance-formatted version of the Stanford Cars fine-grained benchmark — 8,144 photographs across 196 make/model/year classes — sourced from Multimodal-Fatima/StanfordCars_train. Each row carries the inline JPEG bytes, the integer class id, a BLIP-generated caption inherited from the source mirror, and a cosine-normalized CLIP image embedding, all available directly from the Hub at hf://datasets/lance-format/stanford-cars-lance/data.
Key… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/stanford-cars-lance.index-cards-parisian-parliamentarians
Parisian Parliamentarians — Scholarly Prosopography Index Cards
26 index cards (across 5 letter-range items A–D · E–H · J–O · P–R · S–Z) from
the parisianparliamentarians scholarly archive on the Internet Archive — a
prosopographical card index of Parisian parliamentary figures, originally compiled
as a research finding aid.
Each row pairs the full-resolution card scan with the Internet Archive's ABBYY OCR
text and full provenance back to the source IA item.
Source &… See the full description on the dataset page: https://huggingface.co/datasets/biglam/index-cards-parisian-parliamentarians.index-cards-harvard-botany-metropolitan-flora
Card File of the Flora of the Metropolitan Parks (Harvard Botany Libraries, 1894–1895)
4,574 botanical specimen index cards from the Harvard University Botany
Libraries' Card File of the Flora of the Metropolitan Parks, 1894–1895 (bulk),
compiled by Walter Deane (1848–1930). Records flora of the Metropolitan Park
system around Boston — Middlesex Fells Reservation, Blue Hills, Norfolk County,
and adjacent areas — with one card per specimen entry: species, locality,
collection date… See the full description on the dataset page: https://huggingface.co/datasets/biglam/index-cards-harvard-botany-metropolitan-flora.index-cards-peabody-newspaper
Peabody Newspaper Index Cards (Peabody Institute Library, MA)
3,694 typewritten index cards from the Peabody Institute Library — Sutton
Room Local History Resource Center (Peabody, Massachusetts), indexing people,
events, and news in South Danvers / Peabody as recorded in local newspapers.
The information was typed onto cards over decades by library staff as the local
newspaper-of-record archive's principal finding aid.
Plus a companion "Poor Family" genealogy index from the same… See the full description on the dataset page: https://huggingface.co/datasets/biglam/index-cards-peabody-newspaper.index-cards-southborough-vital-records
Southborough Town Clerk — Vital Records & Veteran Index Cards (MA)
8,661 cards from the Southborough (Massachusetts) Town Clerk office, covering:
Death Index Cards, 1850–2015 — the town clerk's running death index covering ~165 years of Southborough deaths, one card per decedent with surname/given-name/date.
Veteran Card Index + Veteran Grave Registration Card Index — companion indices to Southborough's veteran-affairs records, indexing veterans buried in town cemeteries.… See the full description on the dataset page: https://huggingface.co/datasets/biglam/index-cards-southborough-vital-records.index-card-blank-content
Index-card blank / content / divider classifier — dataset
Cropped single archival index cards labelled blank, content, or divider, for
training a tiny CPU pre-filter that skips blank/divider cards before expensive VLM metadata
extraction in card-catalogue digitisation pipelines.
Two collections: Boston Public Library (BPL) FRC shelf-list cards and National Library
of Scotland (NLS) Advocates Library cards. Styles differ, so evaluate per collection.
How it was made… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/index-card-blank-content.index-cards-navy-nurse-corps
US Navy Nurse Corps — Index Cards
25 biographical service cards from the US Navy's Bureau of Medicine and Surgery
(BUMED) History Office, one card per nurse, harvested from their Internet Archive
collection. Includes members of the "Sacred Twenty" — the first female Navy
nurses, appointed in 1908 (Esther V. Hasson, Lena S. Higbee, Elizabeth Hewitt,
Della Knight, M. Estelle Hine, Mary Du Bose, Margaret Murray, Sara Cox, Sara Myer,
Ada M. Pendleton, Florence T. Milburn, Clare L. De… See the full description on the dataset page: https://huggingface.co/datasets/biglam/index-cards-navy-nurse-corps.cardiac_cine_acdc
ACDC (Cardiac Cine-MRI)
ACDC (Automatic Cardiac Diagnosis Challenge, MICCAI 2017) is a cine‑MRI dataset for cardiac segmentation.This repository contains processed NIfTI files in Data/processed_output/acdc format.
Dataset Summary
Modality: Cardiac cine‑MRI (NIfTI)
Task: Segmentation of LV, RV, and myocardium
Frames: ED/ES + full SAX time series (sax_t)
Labels: LV/RV cavities + myocardium
Splits: train, test (as provided in processed output)
Data… See the full description on the dataset page: https://huggingface.co/datasets/horacesu/cardiac_cine_acdc.billys-cardboard-cars-dataset
Billy's Cardboard Cars Dataset
This is the dataset made to train Billy's Cardboard Cars image LoRA model. All images were AI-Generated.
Dataset Information (Shared with the model):
Size: 30 images in a .zip file with .txt captions.
Captioning: Automatically captioned by Civitai.
Billy's Cardboard Cars Dataset © 2025 by Robb-0 is licensed under CC BY 4.0
openm3chest-cardiology-test-200
OpenM3Chest Cardiology Test 200
This dataset is a sampled test subset prepared from UngLong/openm3chest-labels.
Source
Source repo: UngLong/openm3chest-labels
Source split: test
Subsets/tasks: ['CVD_diagnosis', 'CVD_mortality']
Sampling
Number of samples: 200
Sampling mode: primary_ratio
Primary subset: CVD_diagnosis
Positive label: 1
Positive ratio: 0.5
Seed: 42
Columns
keys: CT scan key / series identifier from source dataset.… See the full description on the dataset page: https://huggingface.co/datasets/ChonJohn171105/openm3chest-cardiology-test-200.orthopedic-screw-imagescar-ukraine-synth
Ukraine Synthetic Vehicle Dataset — Toyota Corolla × BMW 3 Series
Фотореалістичні синтетичні зображення седанів Toyota Corolla та BMW 3 Series в українських урбаністичних сценах, згенеровані через OpenAI gpt-image-2. Кожне зображення семплить з сітки 12 українських міст × 7 погодних умов × 5 часів доби × 6 типів камер (dashcam, CCTV, drone, smartphone, action cam, wall-mounted security).
Зображень: 150
Джерело: OpenAI gpt-image-2 (reference-conditioned edits endpoint)… See the full description on the dataset page: https://huggingface.co/datasets/arrmlet/car-ukraine-synth.car-bdd-fine-grained
Fine-Grained Vehicle Detection Dataset (Corolla × BMW 3-Series, BDD100K-derived)
Real-world dashcam frames with fine-grained make annotations on top of
BDD100K's existing car bounding boxes. Identifies which BDD-labeled cars
are specifically Toyota Corolla sedans or BMW 3-Series sedans.
Images: 1410 unique frames · Annotations: 1524
(558 Corolla + 966 BMW 3-Series)
Source imagery: BDD100K dashcam corpus (Berkeley DeepDrive)
Why this dataset
BDD100K labels every car as… See the full description on the dataset page: https://huggingface.co/datasets/arrmlet/car-bdd-fine-grained.CARScenes
CARScenes
CARScenes is an annotation-only dataset and benchmark for structured scene understanding and failure analysis in autonomous driving.
Dataset Summary
Version: carscenes-v1
Records: 5,192
Sources: Argoverse1, Cityscapes, KITTI, nuScenes
Release type: annotations, schema, split manifests, benchmark tooling, Croissant metadata
Public Links
GitHub: https://github.com/Croquembouche/CARScenes
Paper PDF:… See the full description on the dataset page: https://huggingface.co/datasets/williamhe0712/CARScenes.billys_cardtoys_evertything_dataset
Billy's Cardboard Toys Everything in Natural Language - Image dataset
Created with robb-0/Billys-Cardboard-Cars
We have here 40 cute images of what Robbie's model can create. It actually can do so much more beyond cardboards, as any good image LoRA they add dataset to the basemodel, improving new trained ideas which can be used in many other tasks/subjects.
Here those 40 images are captioned in NATURAL LANGUAGE, no booru tags. I decided to do that, because from SDXL onwards… See the full description on the dataset page: https://huggingface.co/datasets/eastenddan/billys_cardtoys_evertything_dataset.
