datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
obelics_seed2_tokensPart of the OBELISC data set, including 32 Million samples, please refer to dataset.py to use this data
ReCo-Data
ReCo-Data Dataset Card
Introduction
ReCo-Data is a large-scale, high-quality video editing dataset comprising 500K+ instruction-video pairs. This card provides its statistics, collection pipeline, and dataset format.
1. Dataset Statistics
Statistics
Figure Caption:
(a) Overview of scale
(b) Task distribution showing balanced quantities: Replace (156.6K), Style (130.6K), Remove (121.6K), and Add (115.6K). Human evaluation on 200 randomly… See the full description on the dataset page: https://huggingface.co/datasets/HiDream-ai/ReCo-Data.S1-MMAlignS1-MMAlign
A Large-Scale Multi-Disciplinary Scientific Multimodal Dataset
S1-MMAlign is a large-scale, multi-disciplinary multimodal dataset comprising over 15.5 million high-quality image-text pairs derived from 2.5 million open-access scientific papers.
Multimodal learning has revolutionized general domain tasks, yet its application in scientific discovery is hindered by the profound semantic gap between complex scientific imagery and sparse textual descriptions. S1-MMAlign aims to… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/S1-MMAlign.captioned-ai-music-snippets
Dataset Overview
A collection of short audio snippets (3–30 seconds) extracted from publicly shared Suno‑generated songs and captioned with Gemini Flash 2.0. Designed specifically to train and evaluate audio captioning models.
Source
Clips are randomly cut from the songs referenced in the nyuuzyou/suno repository.
Captioning
All excerpts have been annotated using Gemini Flash 2.0 for high‑quality, human‑readable audio descriptions.
License
Apache 2.0
wds_fgvc_aircraftPhysical-AI-AV-US
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 2,789,773 samples from 150 000
driving scenes (18 seconds per scene, sampled at 1 Hz) recorded in the
United States.
Format
WebDataset — 100 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.png
Front-facing wide-angle camera frame (640 × 360 px)
{key}.json
Metadata (see schema below)… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-US.OpenSubject
OpenSubject Dataset
OpenSubject is a video-derived large-scale corpus with 2.5M samples and 4.35M images for subject-driven generation and manipulation, as presented in the paper OpenSubject: Leveraging Video-Derived Identity and Diversity Priors for Subject-driven Image Generation and Manipulation.
Project Page & Code
See the main repository for more details and code: OpenSubject
Dataset Structure
OpenSubject/
├── Images_packages/ # Compressed image… See the full description on the dataset page: https://huggingface.co/datasets/AIPeanutman/OpenSubject.ai-music-deduplicated
AI Music Deduplicated
A large-scale collection of AI-generated music from five platforms: Mureka, Riffusion, Sonauto, Suno, and Udio. Each song includes the original audio file and its full platform metadata as a JSON sidecar.
Overview
Subset
Songs
Tar Files
Size
Audio Format
Source Platform
mureka
~312K
49
~981 GB
.mp3
Mureka
riffusion
~105K
14
~266 GB
.m4a
Riffusion
sonauto
~15K
2
~25 GB
.ogg
Sonauto
suno
~307K
65
~1.3 TB
.mp3
Suno
udio~126K
33
~642… See the full description on the dataset page: https://huggingface.co/datasets/ai-music/ai-music-deduplicated.synthetic-cyrillic-largeai-mMemeEffect-382K-audioWe are releasing the audio files that we have collected from MemeEffect-382K dataset. All the files are being shared as .tar files and files are rnamed using their respective id that can be found through the metadata.
We share these files as-part of research initiative.
VenusREMplanartrackplusRU-AI-noise
RU-AI: A Large Multimodal Dataset for Machine Generated Content Detection
This is the noise agumented data for paper: RU-AI: A Large Multimodal Dataset for Machine Generated Content Detection
The original dataset is avaliable at zenodo:
https://zenodo.org/records/11406538
The official repo is avaliable at:
https://github.com/ZhihaoZhang97/RU-AI
Reference
We are appreciated the open-source community for the datasets and the models.
Microsoft COCO: Common Objects in… See the full description on the dataset page: https://huggingface.co/datasets/zzha6204/RU-AI-noise.bucketGUI-AIMA-multiturnparlament_parla_v3
Dataset Card for ParlamentParla v3 - Speech Corpus of Catalan Parliamentary Sessions
A speech corpus composed of Catalan Parliamentary Sessions.The v3 and last version of the corpus includes both clean and other quality segments, divided into short segments (less than 30 seconds) and long segments (more than 30 seconds). The total dataset encompasses 1059h 48m 04s of speech, including 945h 51m 06s for the short segments and 113h 56m 58s for the long segments, with a total of… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/parlament_parla_v3.vel_commons_wikidata
Visual Entity Linking: Wikimedia Commons & Wikidata
This dataset allows to train and evaluate ML models that link Wikimedia Commons images to the Wikidata items they depict.
Disclaimer: All images contained in this dataset are generally assumed to be freely usable (as intended for Wikimedia Commons). Each image's license and author/
uploader is - to the best of our ability - reported in its metadata (see section Dataset Structure). If you want your image's attribution changed or the… See the full description on the dataset page: https://huggingface.co/datasets/aiintelligentsystems/vel_commons_wikidata.Thalia
Thalia: A Global, Multi-Modal Dataset for Volcanic Activity Monitoring
Paper | GitHub | Interactive Demo (Colab)
Thalia is a global, multi-modal dataset for volcanic activity monitoring through Satellite-based Interferometric Synthetic Aperture Radar (InSAR) imagery. Building upon the Hephaestus dataset, Thalia provides higher-resolution, multi-source, and multi-temporal data in a machine-learning-ready format.
Dataset Overview
Thalia consists of 38 spatiotemporal… See the full description on the dataset page: https://huggingface.co/datasets/orion-ai-lab/Thalia.dan-new-webp-trainaislop-videosTransNormal-Synthetic
TransNormal-Synthetic Dataset
Physics-based synthetic dataset for transparent object normal estimation.
This dataset accompanies the paper:
TransNormal: Dense Visual Semantics for Diffusion-based Transparent Object Normal Estimation
Mingwei Li, Hehe Fan, Yi Yang
arXiv:2602.00839 | Project Page | Code
Dataset Description
TransNormal-Synthetic is a physics-based rendered dataset featuring transparent laboratory equipment (beakers, flasks, test tubes, etc.) with… See the full description on the dataset page: https://huggingface.co/datasets/Longxiang-ai/TransNormal-Synthetic.Warwick-STEM
Warwick STEM Dataset (WebDataset)
A collection of 19,769 experimental scanning transmission electron microscopy (STEM) images from the University of Warwick, spanning hundreds of diverse materials projects collected between 2010 and 2018.
Dataset Description
This dataset contains experimental STEM images originally published as part of the Warwick Electron Microscopy Datasets by Jeffrey Ede. The images cover a wide range of materials and imaging conditions, making them… See the full description on the dataset page: https://huggingface.co/datasets/Stemson-AI/Warwick-STEM.1024
MidJourney
NijiJourney
E-Shushu
Compressed by a factor of 64 pixels in JPEG, 97 quality, maximum side length of 1024.
Mixed labeling using different models:
Human prompts in MJ/NJ
Long captions (LLaVA)
Short captions (LLaVA + LLaMA)
Medium captions (Moondream)
AISHELL6-Whisper
🗣️ AISHELL6-Whisper
AISHELL6-Whisper is a large-scale open-source Chinese Mandarin audio-visual whisper speech dataset,containing 30 hours each of whisper and parallel normal speech, with synchronized frontal RGB facial videos.
📘 Dataset Summary
Property
Description
Language
Chinese (Mandarin, ZH)
License
CC BY-NC-SA 4.0
Duration
~60 hours total (30 h whisper + 30 h normal)
Speakers
167 total (121 with RGB-D, 46 audio-only)
Environment
Controlled… See the full description on the dataset page: https://huggingface.co/datasets/SMIIP-lab/AISHELL6-Whisper.physical-ai-bench-artifactspatho-ssl-data-curation
Revisiting Automatic Data Curation for Vision Foundation Models in Digital Pathology
Abstract Vision foundation models (FMs) are accelerating the devel- opment of digital pathology algorithms and transforming biomedical research. These models learn, in a self-supervised manner, to represent histological features in highly heterogeneous tiles extracted from whole-slide images (WSIs) of real-world patient samples. The performance of these FMs is significantly influenced by the size… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/patho-ssl-data-curation.dtasettarai-generated-songswds_fgvc_aircraft_test
FGVC-Aircraft (Test set only)
Original paper: Fine-Grained Visual Classification of Aircraft
Homepage: https://www.robots.ox.ac.uk/~vgg/data/fgvc-aircraft/
Bibtex:
@techreport{maji13fine-grained,
title = {Fine-Grained Visual Classification of Aircraft},
author = {S. Maji and J. Kannala and E. Rahtu
and M. Blaschko and A. Vedaldi},
year = {2013},
archivePrefix = {arXiv},
eprint = {1306.5151},
primaryClass = "cs-cv"… See the full description on the dataset page: https://huggingface.co/datasets/djghosh/wds_fgvc_aircraft_test.
