datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
envsscml-tts
Dataset Card for CML-TTS
Dataset Summary
CML-TTS is a recursive acronym for CML-Multi-Lingual-TTS, a Text-to-Speech (TTS) dataset developed at the Center of Excellence in Artificial Intelligence (CEIA) of the Federal University of Goias (UFG).
CML-TTS is a dataset comprising audiobooks sourced from the public domain books of Project Gutenberg, read by volunteers from the LibriVox project. The dataset includes recordings in Dutch, German, French, Italian, Polish… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/cml-tts.piqaTo apply eyeshadow without a brush, should I use a cotton swab or a toothpick?
Questions requiring this kind of physical commonsense pose a challenge to state-of-the-art
natural language understanding systems. The PIQA dataset introduces the task of physical commonsense reasoning
and a corresponding benchmark dataset Physical Interaction: Question Answering or PIQA.
Physical commonsense knowledge is a major challenge on the road to true AI-completeness,
including robots that interact with the world and understand natural language.
PIQA focuses on everyday situations with a preference for atypical solutions.
The dataset is inspired by instructables.com, which provides users with instructions on how to build, craft,
bake, or manipulate objects using everyday materials.
The underlying task is formualted as multiple choice question answering:
given a question `q` and two possible solutions `s1`, `s2`, a model or
a human must choose the most appropriate solution, of which exactly one is correct.
The dataset is further cleaned of basic artifacts using the AFLite algorithm which is an improvement of
adversarial filtering. The dataset contains 16,000 examples for training, 2,000 for development and 3,000 for testing.yahoo-finance-data
The Financial data from Yahoo!
*** Key Points to Note ***
All financial data is sourced from Yahoo!Ⓡ Finance, Nasdaq!Ⓡ, and the U.S. Department of the Treasury via publicly available APIs, and is intended for research and educational purposes.
I will update the data regularly, and you are welcome to follow this project and use the data.
Each time the data is updated, I will record the update time in spec.json.
Data Usage Instructions
Use DuckDB… See the full description on the dataset page: https://huggingface.co/datasets/defeatbeta/yahoo-finance-data.webshop-data3DCode
Project page
Paper
Code
3dcodebench.com
arXiv:2606.01057
gaoypeng/3dcodebench
News
[06/01/2026] Paper released on arXiv: 3DCodeBench: Benchmarking Agentic Procedural 3D Modeling Via Code.
Note. This is an open-source reproduction of 3DCodeBench.
⚠️ Under final check. The 3DCodeData/ code is still undergoing final
quality review and may contain occasional issues (non-executable scripts, mismatched
captions/renders, or imperfect geometry). If you run… See the full description on the dataset page: https://huggingface.co/datasets/YipengGao/3DCode.mnist
Dataset Card for MNIST
Dataset Summary
The MNIST dataset consists of 70,000 28x28 black-and-white images of handwritten digits extracted from two NIST databases. There are 60,000 images in the training dataset and 10,000 images in the validation dataset, one class per digit so a total of 10 classes, with 7,000 images (6,000 train images and 1,000 test images) per class.
Half of the image were drawn by Census Bureau employees and the other half by high school students… See the full description on the dataset page: https://huggingface.co/datasets/ylecun/mnist.yodas-granary
Dataset Card for YODAS-Granary
Repository: NeMo-speech-data-processor: Granary
Paper: Granary: Speech Recognition and Translation Dataset in 25 European Languages
Shared by: ESPnet
Dataset Description
YODAS-Granary is a curated subset of the larger nvidia/Granary dataset, focusing on high-quality pseudo-labeled speech data for Automatic Speech Recognition (ASR) and Automatic Speech Translation (AST) across 23 European languages.
Overview… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas-granary.L2DTL;DR of L2D, the world's largest self-driving dataset! Read more about L2D on the official Huggingface blog: LeRobot goes to driving school
90+ TeraBytes of multimodal data (5000+ hours of driving) from 30 cities in Germany
6x surrounding HD cameras and complete vehicle state: Speed/Heading/GPS/IMU
Continuous: Gas/Brake/Steering and discrete actions: Gear/Turn Signals
Environment state: Lane count, Road type (highway|residential), Road surface (asphalt, cobbled, sett), Max speed limit.… See the full description on the dataset page: https://huggingface.co/datasets/yaak-ai/L2D.yodasUpdates
2024/07/09: we also uploaded a new version of YODAS as YODAS2, it provides unsegmented audios and higher sampling rate (24k)
README
This is the YODAS manual/automatic subset from our YODAS dataset, it has 369,510 hours of speech.
This dataset contains audio utterances and corresponding captions (manual or automatic) from YouTube. Note that manual caption only indicates that it is uploaded by users, but not necessarily transcribed by a human
For more details about YODAS… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas.yodas2YODAS2 is the long-form dataset from YODAS dataset.
It provides the same dataset as espnet/yodas but YODAS2 has the following new features:
formatted in the long-form (video-level) where audios are not segmented.
audios are encoded using higher sampling rates (i.e. 24k)
For detailed information about YODAS dataset, please refer to our paper and the espnet/yodas repo.
Usage:
Each data point corresponds to an entire video on YouTube, it contains the following fields:
video_id:… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas2.yenle84622checkpoint
Dataset Card for LLaVA-Video-178K
Uses
This dataset is used for the training of the LLaVA-Video model. We only allow the use of this dataset for academic research and education purpose. For OpenAI GPT-4 generated data, we recommend the users to check the OpenAI Usage Policy.
Data Sources
For the training of LLaVA-Video, we utilized video-language data from five primary sources:
LLaVA-Video-178K: This dataset includes 178,510 caption entries, 960,792 open-ended… See the full description on the dataset page: https://huggingface.co/datasets/YYF111/checkpoint.DimeggLIBERO-datasets
LIBERO Datasets
This is a repo that stores the LIBERO datasets. The structure of the dataset can be found below:
libero_object/
libero_spatial/
libero_goal/
libero_90/
libero_10/
Demonstrations of each task is stored in a hdf5 file. Please refer to download script from the official LIBERO repo for more details.
aime_2025
AIME 2025
This dataset contains 30 problems from the 2025 AIME tests, including:
AIME I: 15 problems
AIME II: 15 problems
TexVerse
TexVerse: A Universe of 3D Objects with High-Resolution Textures
Yibo Zhang1,2, Li Zhang1,3, Rui Ma2 *, Nan Cao1,4
1Shanghai Innovation Institute
2Jilin University
3Fudan University
4Tongji University
* Corresponding Author
TexVerse is a large-scale 3D dataset featuring high-resolution textures. Its key characteristics include:
Scale & Source: TexVerse dataset has 858,669 unique 3D models curated from… See the full description on the dataset page: https://huggingface.co/datasets/YiboZhang2001/TexVerse.oi-devyodas2_sidon
YODAS2-Sidon
Overview
This dataset is a cleansed version of YODAS-2 with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling.
YODAS-2 is a massive, multilingual YouTube-derived dataset. We have applied the Sidon restoration model to remove background noise and enhance audio quality, making it suitable for high-quality generation tasks.
We resampled original sidon output to 24kHz due to a storage constraints.
The dataset is provided in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/yodas2_sidon.RGB-Event-ISP-DatasetTexVerse-1KPeRFception-v1-2Slides-Align
Slides-Align: Human Preference Rankings for AI-Generated Presentations
Project Page | Paper | GitHub
Overview
Slides-Align is a human preference dataset for evaluating AI-generated slide presentations, introduced as part of the SlidesGen-Bench framework. It contains 1,326 human rankings comparing presentations generated by 9 different AI slide generation products across 7 scenario categories and 187 unique topics.
This dataset enables:
🎯 Benchmarking AI slide… See the full description on the dataset page: https://huggingface.co/datasets/Yqy6/Slides-Align.World-SimReady-Home
WorldSimReady-Home
Dataset description
CAD-based SimReady assets
Optimized CAD assets with configured collision and physical properties.
Manually reviewed scenes
Physics configuration reviewed for every household scene.
Scalable task generation
Batch simulation data across robot embodiments and tasks.
WorldSimReady-Home is built from CAD-based object assets, optimized and enriched with physical properties to create simulation-ready assets. The… See the full description on the dataset page: https://huggingface.co/datasets/Yootta/World-SimReady-Home.alpaca-cleaned
Dataset Card for Alpaca-Cleaned
Repository: https://github.com/gururise/AlpacaDataCleaned
Dataset Description
This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset:
Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused GPT3 to hallucinate an answer.
"instruction":"Summarize… See the full description on the dataset page: https://huggingface.co/datasets/yahma/alpaca-cleaned.code-world-model-project-page-videos
Code World Model Project Page Videos
Public research-demo video assets used by the Code World Model project page.
The gallery/ directory contains aligned RGB and proxy videos for interactive comparison.
MASIV
MASIV Multi-Sequence Dataset
Toward Material-Agnostic System Identification from Videos
ICCV 2025
Yizhou Zhao1, Haoyu Chen1, Chunjiang Liu1, Zhenyang Li2, Charles Herrmann3, Junhwa Hur3, Yinxiao Li3, Ming‑Hsuan Yang4, Bhiksha Raj1, Min Xu1*
1Carnegie Mellon University 2University of Alabama at Birmingham 3Google 4UC Merced
Introduction
The MASIV Multi-Sequence Dataset is a synthetic dataset generated by Genesis to evaluate the generalization of data-driven… See the full description on the dataset page: https://huggingface.co/datasets/yizhouz/MASIV.yelp_review_full
Dataset Card for YelpReviewFull
Dataset Summary
The Yelp reviews dataset consists of reviews from Yelp.
It is extracted from the Yelp Dataset Challenge 2015 data.
Supported Tasks and Leaderboards
text-classification, sentiment-classification: The dataset is mainly used for text classification: given the text, predict the sentiment.
Languages
The reviews were mainly written in english.
Dataset Structure
Data Instances
A… See the full description on the dataset page: https://huggingface.co/datasets/Yelp/yelp_review_full.sd3_5_fine_sixcard
DreamBooth training example
DreamBooth is a method to personalize text2image models like stable diffusion given just a few(3~5) images of a subject.
The train_dreambooth.py script shows how to implement the training procedure and adapt it for stable diffusion.
Running locally with PyTorch
Installing the dependencies
Before running the scripts, make sure to install the library's training dependencies:
Important
To make sure you can successfully run the latest… See the full description on the dataset page: https://huggingface.co/datasets/yyyzzzzyyy/sd3_5_fine_sixcard.dataset-with-standalone-yamlThis is a test dataset used in the datasets library CI
