datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RenderedTextThis dataset has been created by Stability AI and LAION.
This dataset contains 12 million 1024x1024 images of handwritten text written on a digital 3D sheet of paper generated using Blender geometry nodes and rendered using Blender Cycles. The text has varying font size, color, and rotation, and the paper was rendered under random lighting conditions.
Note that, the first 10 million examples are in the root folder of this dataset repository and the remaining 2 million are in ./remaining (due… See the full description on the dataset page: https://huggingface.co/datasets/wendlerc/RenderedText.wds_imagenet-rrobotwin2.0-fastwam
RobotWin 2.0 (Preprocessed LeRobot v2.1 Release)
This repository releases our preprocessed RoboTwin / RobotWin 2.0 dataset in LeRobot v2.1 format for the open-source release of Fast-WAM: Do World Action Models Need Test-time Future Imagination?
This is not the official upstream RoboTwin release. It is our paper-specific processed version prepared to support training, evaluation, and reproducibility for our project.
This Hugging Face repository distributes the dataset as split… See the full description on the dataset page: https://huggingface.co/datasets/yuanty/robotwin2.0-fastwam.Emilia-YODAS-ENrobopoint-data
RoboPoint Dataset Card
Dataset details
This dataset contains 1432K image-QA instances used to fine-tune RoboPoint, a VLM for spatial affordance prediction. It consists of the following parts:
347K object reference instances from a synthetic data pipeline;
320K free space reference instances from a synthetic data pipeline;
100K object detection instaces from LVIS;
150K GPT-generated instruction-following instances from liuhaotian/LLaVA-Instruct-150K;
515K general-purpose… See the full description on the dataset page: https://huggingface.co/datasets/wentao-yuan/robopoint-data.ReCo-Data
ReCo-Data Dataset Card
Introduction
ReCo-Data is a large-scale, high-quality video editing dataset comprising 500K+ instruction-video pairs. This card provides its statistics, collection pipeline, and dataset format.
1. Dataset Statistics
Statistics
Figure Caption:
(a) Overview of scale
(b) Task distribution showing balanced quantities: Replace (156.6K), Style (130.6K), Remove (121.6K), and Add (115.6K). Human evaluation on 200 randomly… See the full description on the dataset page: https://huggingface.co/datasets/HiDream-ai/ReCo-Data.reazonspeechlaions_got_talent_rawmelee-ranked-replays
Melee Ranked Replays
Anonymized Slippi ranked replays (platinum+) from Super Smash Bros. Melee,
sharded by character and rank pair. Built for behavior-cloning and other
replay-driven ML work on Melee — notably MIMIC.
Contents
Raw .slp files grouped into tarballs by (character, rank_pair, source_archive),
organized into per-character folders:
{CHAR}/
{CHAR}_{rank_pair}_a{N}.tar.gz
metadata/
metadata_a{N}.json
Characters (25): BOWSER, CPTFALCON, DK, DOC, FALCO… See the full description on the dataset page: https://huggingface.co/datasets/erickfm/melee-ranked-replays.RevealLayer-100K
RevealLayer Open Dataset
RevealLayer Open is the open-source dataset accompanying RevealLayer: Disentangling Hidden and Visible Layers via Occlusion-Aware Image Decomposition.
Paper: https://arxiv.org/html/2605.11818v1 Accepted by ICML 2026
RevealLayer studies box-guided layered image decomposition for natural images. Given an RGB image and instance bounding boxes, the task is to decompose the scene into a clean background and object-level foreground layers, where each… See the full description on the dataset page: https://huggingface.co/datasets/qihoo360/RevealLayer-100K.cc12m-recaptionedraw_primitive_datasets
Datacard
This is the official fine-tuning dataset provided by VLABench (raw data), with 500 episodes each task. The current version includes 10 primitive tasks.
Source
Project Page: https://vlabench.github.io/
Arxiv Paper: https://arxiv.org/abs/2412.18194
Code: https://github.com/OpenMOSS/VLABench
Uses
Download all archive files and use the following command to extract:
cat vlabench_primitive.tar.gz.* | tar -xzvf -
In the resulting… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/raw_primitive_datasets.Molmo2-ER-RoboPoint
Molmo2-ER · wentao-yuan/robopoint-data
1.43M robotics affordance instruction-tuning examples (pointing + detection + VQA).
This is a re-hosted, loader-ready subset of the upstream dataset, used to train allenai/Molmo2-ER-4B. Files mirror the upstream layout; nothing in the data has been modified.
Upstream source
Original dataset: wentao-yuan/robopoint-data
Paper: RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics (arXiv:2406.10721)
License:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Molmo2-ER-RoboPoint.Emilia-ENBrushDatacc12m-wds-coco-recaptioned
CC12M WebDataset with COCO-style Recaptions
A large-scale image-text dataset containing 3 million images from Conceptual Captions 12M (CC12M) with COCO-style factual descriptions generated using NVIDIA Nemotron Nano 12B v2 VL.
Dataset Overview
Base Dataset: pixparse/cc12m-wds - Conceptual Captions 12M (CC12M)
Images: 3,000,000+ high-quality internet images
Recaption Model: NVIDIA Nemotron Nano 12B v2 VL
Recaption Style: COCO-style factual descriptions (20 words average)… See the full description on the dataset page: https://huggingface.co/datasets/undefined443/cc12m-wds-coco-recaptioned.RH20T
RH20T 640x360 Depth Archives
Unofficial transfer mirror of the RH20T 640x360 depth archives.
Original project: https://rh20t.github.io/
Files are split into 20 GiB parts. Reconstruct with:
cat depth/RH20T_cfgN_depth.tar.gz.part-* > RH20T_cfgN_depth.tar.gz
RH20T uses mixed scene-level licenses: scenes 1-5 are CC BY-SA 4.0;
scenes 6-10 are CC BY-NC 4.0. The data may contain privacy-sensitive
human recordings. Follow the original project's license and privacy notice.
biggest-ru-bookA bigger version of its5Q/bigger-ru-book, the smaller set being a subset of this one. Almost 1000 hours of high-quality audio.
SPI-2M
SPI-2M
We introduce Stylized Pathology Images SPI-2M for stain normalisation via neural style transfer in histopathology.
For full details on dataset sourcing, creation etc please see our paper
Dataset download
The data repo of this repository is organised as follows:
sources: contains the 4096 curated source images zipped together
targets: contains the 512 target images zipped together
stylized: contains 512 .npy files, each has the same index as a corresponding target… See the full description on the dataset page: https://huggingface.co/datasets/R-J/SPI-2M.redcaps5m_resizedRaCig-Dataosu-beatmaps
osu! Beatmaps Dataset (WebDataset)
A collection of ranked/loved osu! beatmaps with audio and chart data, in WebDataset format.
Dataset Variants
Variant
Audio Format
Description
original
MP3/OGG/WAV
Full quality original audio files
compressed
64kbps Mono Opus
Compressed audio for smaller download
from datasets import load_dataset
# Load original audio variant
ds = load_dataset("project-riz/osu-beatmaps", "original", streaming=True)
# Load compressed… See the full description on the dataset page: https://huggingface.co/datasets/project-riz/osu-beatmaps.objaverse_rendering_setwikipedia_rureazon-speech-v2-clone
Reazon Speech v2 dataset mirror
Original Dataset Source
Hugging Face Dataset Page: reazon-research/reazonspeech
Project Page: Reazon Research
License
This dataset is a mirror of the original Reazon Speech v2 dataset, but on 🤗 server (so may be faster). This dataset is licensed under the CDLA-Sharing-1.0. The original dataset comes with the following restriction:
TO USE THIS DATASET, YOU MUST AGREE THAT YOU WILL USE THE DATASET SOLELY FOR THE PURPOSE OF… See the full description on the dataset page: https://huggingface.co/datasets/litagin/reazon-speech-v2-clone.unsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_1ImageNet-RenditiontestdataObjaverseXL_github_rendersmagicdata_ramc
