datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Physical-AI-AV-US
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 2,789,773 samples from 150 000
driving scenes (18 seconds per scene, sampled at 1 Hz) recorded in the
United States.
Format
WebDataset — 100 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.png
Front-facing wide-angle camera frame (640 × 360 px)
{key}.json
Metadata (see schema below)… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-US.S1-MMAlignS1-MMAlign
A Large-Scale Multi-Disciplinary Scientific Multimodal Dataset
S1-MMAlign is a large-scale, multi-disciplinary multimodal dataset comprising over 15.5 million high-quality image-text pairs derived from 2.5 million open-access scientific papers.
Multimodal learning has revolutionized general domain tasks, yet its application in scientific discovery is hindered by the profound semantic gap between complex scientific imagery and sparse textual descriptions. S1-MMAlign aims to… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/S1-MMAlign.wds_fgvc_aircraftsynthetic-cyrillic-largeOpenSubject
OpenSubject Dataset
OpenSubject is a video-derived large-scale corpus with 2.5M samples and 4.35M images for subject-driven generation and manipulation, as presented in the paper OpenSubject: Leveraging Video-Derived Identity and Diversity Priors for Subject-driven Image Generation and Manipulation.
Project Page & Code
See the main repository for more details and code: OpenSubject
Dataset Structure
OpenSubject/
├── Images_packages/ # Compressed image… See the full description on the dataset page: https://huggingface.co/datasets/AIPeanutman/OpenSubject.GUI-AIMA-multiturnvel_commons_wikidata
Visual Entity Linking: Wikimedia Commons & Wikidata
This dataset allows to train and evaluate ML models that link Wikimedia Commons images to the Wikidata items they depict.
Disclaimer: All images contained in this dataset are generally assumed to be freely usable (as intended for Wikimedia Commons). Each image's license and author/
uploader is - to the best of our ability - reported in its metadata (see section Dataset Structure). If you want your image's attribution changed or the… See the full description on the dataset page: https://huggingface.co/datasets/aiintelligentsystems/vel_commons_wikidata.dan-new-webp-trainTransNormal-Synthetic
TransNormal-Synthetic Dataset
Physics-based synthetic dataset for transparent object normal estimation.
This dataset accompanies the paper:
TransNormal: Dense Visual Semantics for Diffusion-based Transparent Object Normal Estimation
Mingwei Li, Hehe Fan, Yi Yang
arXiv:2602.00839 | Project Page | Code
Dataset Description
TransNormal-Synthetic is a physics-based rendered dataset featuring transparent laboratory equipment (beakers, flasks, test tubes, etc.) with… See the full description on the dataset page: https://huggingface.co/datasets/Longxiang-ai/TransNormal-Synthetic.1024
MidJourney
NijiJourney
E-Shushu
Compressed by a factor of 64 pixels in JPEG, 97 quality, maximum side length of 1024.
Mixed labeling using different models:
Human prompts in MJ/NJ
Long captions (LLaVA)
Short captions (LLaVA + LLaMA)
Medium captions (Moondream)
Warwick-STEM
Warwick STEM Dataset (WebDataset)
A collection of 19,769 experimental scanning transmission electron microscopy (STEM) images from the University of Warwick, spanning hundreds of diverse materials projects collected between 2010 and 2018.
Dataset Description
This dataset contains experimental STEM images originally published as part of the Warwick Electron Microscopy Datasets by Jeffrey Ede. The images cover a wide range of materials and imaging conditions, making them… See the full description on the dataset page: https://huggingface.co/datasets/Stemson-AI/Warwick-STEM.patho-ssl-data-curation
Revisiting Automatic Data Curation for Vision Foundation Models in Digital Pathology
Abstract Vision foundation models (FMs) are accelerating the devel- opment of digital pathology algorithms and transforming biomedical research. These models learn, in a self-supervised manner, to represent histological features in highly heterogeneous tiles extracted from whole-slide images (WSIs) of real-world patient samples. The performance of these FMs is significantly influenced by the size… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/patho-ssl-data-curation.dtasettarZOD-Mini-2D-Road-Scenes
ZOD-Mini-2D-Road-Scenes
The ZOD-Mini-2D-Road-Scenes dataset is derived from the Zenseact Open Dataset (ZOD), property of Zenseact AB (© 2022 Zenseact AB), and is licensed under the permissive CC BY-SA 4.0. Any public use, distribution, or display of this dataset must contain this entire notice:
For this dataset, Zenseact AB has taken all reasonable measures to remove all personally identifiable information, including faces and license plates. To the extent that you like to request… See the full description on the dataset page: https://huggingface.co/datasets/8bits-ai/ZOD-Mini-2D-Road-Scenes.wds_fgvc_aircraft_test
FGVC-Aircraft (Test set only)
Original paper: Fine-Grained Visual Classification of Aircraft
Homepage: https://www.robots.ox.ac.uk/~vgg/data/fgvc-aircraft/
Bibtex:
@techreport{maji13fine-grained,
title = {Fine-Grained Visual Classification of Aircraft},
author = {S. Maji and J. Kannala and E. Rahtu
and M. Blaschko and A. Vedaldi},
year = {2013},
archivePrefix = {arXiv},
eprint = {1306.5151},
primaryClass = "cs-cv"… See the full description on the dataset page: https://huggingface.co/datasets/djghosh/wds_fgvc_aircraft_test.dtasettar23PhD-webdataset
PhD Webdataset
This repository contains the packaged version of PhD. For a detailed introduction to PhD, please visit the official website.
Overview
The PhD Webdataset is designed to facilitate easy access and usage of the PhD dataset. It includes various fields in 'json' key. The data in this repo is totally the same as in PhD.
Installation
Ensure you have Hugging Face's datasets library installed. You can install it via pip:
pip install datasets… See the full description on the dataset page: https://huggingface.co/datasets/AIMClab-RUC/PhD-webdataset.DicFace-test_datasetaigcAsian photography dataset
win3000: about 18k asian celebrity photo.
jiepaigou: streetsnap and celebrity
cybesx: about 13k street photography
jelapang-padimjnj_tarPhysical-AI-AV-FR
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 29,909 samples from 150 000
driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the
United States.
Format
WebDataset — 3 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.jpg
Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px)
{key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-FR.mjnj640aitw_imagePhysical-AI-AV-ES
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 29,674 samples from 150 000
driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the
United States.
Format
WebDataset — 3 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.jpg
Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px)
{key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-ES.ai-check-20m-plus
AI Check Dataset Plus (20M)
This is the webdataset subset dataset for AI checking.
Images here are resized to min(width, height) <= 640.
How to Use It
from datasets import load_dataset
dataset = load_dataset('deepghs/ai-check-20m-plus')
print(dataset["train"][0])
Images
20000000 images in total.
Split
Image Count
Total Size
train
18823238
981 GB
test
594988
31 GB
val
581774
30.4 GB
Class
Image Count
Total Size
ai
9977437
555 GB… See the full description on the dataset page: https://huggingface.co/datasets/just-a-try/ai-check-20m-plus.AIWD16-TextPhysical-AI-AV-DE
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 324,105 samples from 150 000
driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the
United States.
Format
WebDataset — 10 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.jpg
Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px)
{key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-DE.wds_fgvc_aircraftPhysical-AI-AV-IT
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 29,991 samples from 150 000
driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the
United States.
Format
WebDataset — 3 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.jpg
Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px)
{key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-IT.
