datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stanford_cars
Stanford Cars Dataset
Dataset Overview
Splits:
Training: 8144 images used for model training.
Test: 8041 images used for evaluation.
Contrast: 8041 images with high contrast for robustness testing.
Gaussian Noise: 8041 images corrupted by Gaussian noise for robustness testing.
Impulse Noise: 8041 images corrupted by impulse noise for robustness testing.
JPEG Compression: 8041 compressed images for robustness testing.
Motion Blur: 8041 images with motion blur for… See the full description on the dataset page: https://huggingface.co/datasets/tanganke/stanford_cars.CarlaOcc
Database_structure
CarlaOcc/
├── CarlaOccV1/
│ ├── calib/
│ │ └── calib.yaml
│ ├── splits/
│ │ ├── test.txt
│ │ ├── train.txt
│ │ └── val.txt
│ ├── SceneMeshes/
│ │ ├── fg_actors/
│ │ ├── fg_actor_occ/
│ │ └── TownXX_Opt/
│ │ ├── bg_actors/
│ │ └── bg_actor_occ/
│ ├── TownXX_Opt_SeqXX/
│ │ ├── poses/
│ │ │ ├── cam_00.txt
│ │ │ └── lidar.txt
│ │ ├── rgb/
│ │ │ ├── image_00/
│ │ │ │ ├── 0000.png… See the full description on the dataset page: https://huggingface.co/datasets/fengyi233/CarlaOcc.car-damage-dataset
Car Damage Images
A raw image collection for vehicle damage assessment. Unlabeled: these images
have no annotations yet and are intended as source material for labeling or
pre-training.
Structure
images/
001/ images_001.jpg images_002.jpg ...
002/ ...
... 683 folders
thumbnails/
001/ thumbnail_001.jpg thumbnail_002.jpg ...
... 683 folders
Branch
Files
Folders
Size… See the full description on the dataset page: https://huggingface.co/datasets/Naiscorp/car-damage-dataset.sports-cards
Digital Card Magazine Dataset
This dataset contains sports card images and their associated metadata for training machine learning models in card recognition, text extraction, and value estimation.
Dataset Description
Dataset Summary
A comprehensive collection of sports card images and metadata, including:
Front and back card images
OCR-extracted text with confidence scores
AI-analyzed card attributes
Card details (player, team, year, etc.)
Vision API labels… See the full description on the dataset page: https://huggingface.co/datasets/GotThatData/sports-cards.Stanford-Cars
Dataset Card for "Stanford-Cars"
This is a non-official Stanford-Cars dataset for fine-grained Image Classification.
If you want to download the official dataset, please refer to the here.
carla-autopilot-multimodal-dataset
CARLA Autopilot Multimodal Dataset
This dataset contains synchronized multimodal driving data collected in the CARLA simulator using the autopilot feature. It provides RGB images from multiple cameras, semantic segmentation, LiDAR point clouds, 2D bounding boxes, and ego-vehicle state/control signals across varied weather, maps, and traffic densities.
The dataset is designed for research in autonomous driving, sensor fusion, imitation learning, and self-driving evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/immanuelpeter/carla-autopilot-multimodal-dataset.car-images
Dataset Card for Car Images Dataset
Dataset Description
This dataset contains a collection of car images designed to support tasks such as image classification, car detection, and autonomous driving research. The images feature cars in various settings, including streets and other environments, captured to provide diverse training data for computer vision models.
The dataset aims to enable the development and benchmarking of models that can:
Identify… See the full description on the dataset page: https://huggingface.co/datasets/saifurrehman1234567/car-images.cardio-mark
CardioMark Review Subset
This repository contains an anonymized review subset of the CardioMark benchmark introduced for automated vertebral heart score (VHS) estimation in canine thoracic radiographs.
The subset is provided to support reproducibility and data-quality inspection during peer review.
Dataset Overview
CardioMark is a large-scale benchmark for evaluating the complete VHS measurement pipeline, including:
cardiac landmark localization
geometric VHS estimation… See the full description on the dataset page: https://huggingface.co/datasets/gen-ai-researcher/cardio-mark.cars-object-tracking
Cars Object Tracking
Dataset comprises 10,000+ video frames featuring both light vehicles (cars) and heavy vehicles (minivans). This extensive collection is meticulously designed for research in multi-object tracking and object detection, providing a robust foundation for developing and evaluating various tracking algorithms for road safety system development.
By utilizing this dataset, researchers can significantly enhance their understanding of vehicle dynamics and improve… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/cars-object-tracking.car-images
Dataset Card for Car Images Dataset
Dataset Description
This dataset contains a collection of car images designed to support tasks such as image classification, car detection, and autonomous driving research. The images feature cars in various settings, including streets and other environments, captured to provide diverse training data for computer vision models.
The dataset aims to enable the development and benchmarking of models that can:
Identify different types… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/car-images.synthetic-greeting-cards
Synthetic Greeting Cards Dataset (Synth-GCD)
A modern, fully open synthetic alternative to the proprietary Greeting Cards Dataset (GCD) described in the paper"Weakly Supervised Annotations for Multi-modal Greeting Cards Dataset".
This dataset contains high-quality AI-generated greeting card illustrations with corresponding short messages, designed for research in multimodal classification, retrieval, and generation.
Dataset Summary
Property
Value
Total… See the full description on the dataset page: https://huggingface.co/datasets/gauravs101/synthetic-greeting-cards.toy-car-annotation-YOLOHey everyone,
In my final year project, I created Smart Traffic Management System.The project was to manage traffic lights' delays based on the number of vehicles on road.I made everything worked using Raspberry Pi and pre-recorded videos but it was a "final year project", it was needed to be tested by changing videos frequently which was a kind of hustle. Collecting tons of videos and loading them in Pi was not too hard but it would have cost time, by every time changing names of videos in… See the full description on the dataset page: https://huggingface.co/datasets/tubasid/toy-car-annotation-YOLO.comprehensive-car-damage
Car Front and Rear Damage Detection Dataset
Dataset Summary
This dataset is designed for training and evaluating machine learning models for car damage detection, specifically focusing on front and rear vehicle damages.
It includes high-quality labeled images categorized into six distinct classes:
R_Normal: Rear view of undamaged cars
R_Crushed: Rear view of cars with crushed damage
R_Breakage: Rear view of cars with visible breakage
F_Normal: Front view of… See the full description on the dataset page: https://huggingface.co/datasets/DrBimmer/comprehensive-car-damage.2D_Video_Game_Cartoon_Character_Sprite-Sheets
Dataset Card for Dataset Name
Dataset Details
Experimental composition of 76 cartoon art-style video game character spritesheets. Resized to 512x512, mixed variation of animation styles.
Dataset Description
All images editted using Tiled image editting software as most assets are typically downloaded individually and not in sequence. I compiled each animation sequence into one img to display animations frame-by-frame evenly distributed across some common… See the full description on the dataset page: https://huggingface.co/datasets/mgane/2D_Video_Game_Cartoon_Character_Sprite-Sheets.lalafo-kg-cars
lalafo.kg — Kyrgyzstan Cars (used-car market)
Scraped from lalafo.kg, the largest informal classifieds
board in Kyrgyzstan — messier and larger than the curated boards, and closer to the
real street-level market. Field names are English; values are kept in the original
language (Russian).
Subsets
subset
rows
description
listings
54,518
one row per advertisement (every category, deal and region in scope)
users
49,685
sellers (ad authors), with… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/lalafo-kg-cars.comprehensive-car-damage
Car Front and Rear Damage Detection Dataset
Dataset Summary
This dataset is designed for training and evaluating machine learning models for car damage detection, specifically focusing on front and rear vehicle damages.
It includes high-quality labeled images categorized into six distinct classes:
R_Normal: Rear view of undamaged cars
R_Crushed: Rear view of cars with crushed damage
R_Breakage: Rear view of cars with visible breakage
F_Normal: Front view of… See the full description on the dataset page: https://huggingface.co/datasets/SaiVaibhavS/comprehensive-car-damage.cardboard-box-anomaly-detection
📦 Cardboard Box Anomaly Detection
Descrição
Este é um dataset focado na detecção de anomalias em caixas de papelão. Ele é composto por 553 imagens capturadas de 43 caixas distintas (13 consideradas normais e 30 anômalas).
As imagens foram coletadas em múltiplos ambientes (chão e duas esteiras diferentes), utilizando múltiplos ângulos e duas câmeras de celular diferentes.
Exemplos do Dataset
Caixas Normais (good)
Chão (loc-chao)
Esteira 1… See the full description on the dataset page: https://huggingface.co/datasets/Gabriel8/cardboard-box-anomaly-detection.playing-cardsstandford_cars_masks
Stanford Cars with SAM3 Segmentation Masks
This dataset was created using segmentationAPI.com
A version of the Stanford Cars dataset
augmented with per-image segmentation masks generated by SAM3 (Segment Anything Model 3).
Schema
Column
Type
Description
image
Image
Original car photograph (JPEG)
label
int
Class label (0–195, 196 car models)
split
string
train or test
mask
Image
Binary segmentation mask (grayscale PNG)
bbox
list[int]
Bounding box [x1, y1… See the full description on the dataset page: https://huggingface.co/datasets/segmentationAPIs/standford_cars_masks.pokemon_card_image_for_authenticity_classification
Pokemon Card Image for Authenticity Classification
This dataset contains front/back images of Pokemon cards for authenticity experiments.
Dataset structure
Images/: all image files (.jpeg)
Images/metadata.jsonl: metadata used by Hugging Face imagefolder
labels.csv: flat label file with the same rows as metadata
Columns
image: image object loaded from file
id: image filename (unique id)
side: card side (0 = front, 1 = back)
labels: authenticity label (1 =… See the full description on the dataset page: https://huggingface.co/datasets/stevelohwc/pokemon_card_image_for_authenticity_classification.carambola_disease_classification
Carambola Disease Classification
A dataset for disease classification of Carambola fruits and leaves. The dataset contains raw and augmented versions.The raw dataset contains 2,618 images.Images per class:
Healthy Fruits: 485
Healthy Leaves: 658
Insect Hole leaves: 518
Unhealthy Fruits: 478
Yellow Leaves: 479
The augmented dataset contains 15,000 images.Images per class:
Healthy Fruits: 3,000
Healthy Leaves: 3,000
Insect Hole leaves: 3,000
Unhealthy Fruits: 3,000
Yellow… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/carambola_disease_classification.curated-mnist
Dataset Card for curated-mnist
This is a FiftyOne dataset with 70000 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("CarloColumbo/curated-mnist")
# Launch the App
session = fo.launch_app(dataset)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/CarloColumbo/curated-mnist.index-cards-harvard-botany-metropolitan-flora
Card File of the Flora of the Metropolitan Parks (Harvard Botany Libraries, 1894–1895)
4,574 botanical specimen index cards from the Harvard University Botany
Libraries' Card File of the Flora of the Metropolitan Parks, 1894–1895 (bulk),
compiled by Walter Deane (1848–1930). Records flora of the Metropolitan Park
system around Boston — Middlesex Fells Reservation, Blue Hills, Norfolk County,
and adjacent areas — with one card per specimen entry: species, locality,
collection date… See the full description on the dataset page: https://huggingface.co/datasets/biglam/index-cards-harvard-botany-metropolitan-flora.index-cards-cas-galapagos-stewart-specimens
Alban Stewart's Galápagos Expedition Specimen Cards (CAS Archives, 1905–1906)
1,059 specimen index cards from the California Academy of Sciences Archives,
documenting Alban Stewart's botanical specimens collected on the 1905–1906
California Academy of Sciences Galápagos Expedition. One card per specimen with
species, locality on the islands, collection date and field notes; companion to
Stewart's expedition journal.
Harvested from the 6 IA items
csfa788562626 +
csfa788562626Alpha +… See the full description on the dataset page: https://huggingface.co/datasets/biglam/index-cards-cas-galapagos-stewart-specimens.index-cards-parisian-parliamentarians
Parisian Parliamentarians — Scholarly Prosopography Index Cards
26 index cards (across 5 letter-range items A–D · E–H · J–O · P–R · S–Z) from
the parisianparliamentarians scholarly archive on the Internet Archive — a
prosopographical card index of Parisian parliamentary figures, originally compiled
as a research finding aid.
Each row pairs the full-resolution card scan with the Internet Archive's ABBYY OCR
text and full provenance back to the source IA item.
Source &… See the full description on the dataset page: https://huggingface.co/datasets/biglam/index-cards-parisian-parliamentarians.Wasp_moth_mimicry
Dataset Card for Wasp-Moth Mimicry
This dataset comprises images of pinned insects from four localities around Panama. This represents sympatric species of Euchromiina moths,
Hymenoptera and some Diptera. This dataset was created to understand how similar the phenotype of Euchoemiina moths and the Hymenopteran sympatric species.
Euchromiina moths are characterized by being mimics of different groups of wasps and bees, but to date, no study has quantified their similarity.
This… See the full description on the dataset page: https://huggingface.co/datasets/Sol-Carolina/Wasp_moth_mimicry.car-images
Dataset Card for Car Images Dataset
Dataset Description
This dataset contains a collection of car images designed to support tasks such as image classification, car detection, and autonomous driving research. The images feature cars in various settings, including streets and other environments, captured to provide diverse training data for computer vision models.
The dataset aims to enable the development and benchmarking of models that can:
Identify different types… See the full description on the dataset page: https://huggingface.co/datasets/Smily6820/car-images.stanford-cars-lance
Stanford Cars (Lance Format)
A Lance-formatted version of the Stanford Cars fine-grained benchmark — 8,144 photographs across 196 make/model/year classes — sourced from Multimodal-Fatima/StanfordCars_train. Each row carries the inline JPEG bytes, the integer class id, a BLIP-generated caption inherited from the source mirror, and a cosine-normalized CLIP image embedding, all available directly from the Hub at hf://datasets/lance-format/stanford-cars-lance/data.
Key… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/stanford-cars-lance.cardd_workshop_post_03
Dataset Card for harpreetsahota/cardd_workshop_post_03
This is a FiftyOne dataset with 2816 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("harpreetsahota/cardd_workshop_post_03")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/cardd_workshop_post_03.index-cards-peabody-newspaper
Peabody Newspaper Index Cards (Peabody Institute Library, MA)
3,694 typewritten index cards from the Peabody Institute Library — Sutton
Room Local History Resource Center (Peabody, Massachusetts), indexing people,
events, and news in South Danvers / Peabody as recorded in local newspapers.
The information was typed onto cards over decades by library staff as the local
newspaper-of-record archive's principal finding aid.
Plus a companion "Poor Family" genealogy index from the same… See the full description on the dataset page: https://huggingface.co/datasets/biglam/index-cards-peabody-newspaper.
