datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mbpp
Dataset Card for Mostly Basic Python Problems (mbpp)
Dataset Summary
The benchmark consists of around 1,000 crowd-sourced Python programming problems, designed to be solvable by entry level programmers, covering programming fundamentals, standard library functionality, and so on. Each problem consists of a task description, code solution and 3 automated test cases. As described in the paper, a subset of the data has been hand-verified by us.
Released here as part of… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/mbpp.jat-dataset-tokenized
Dataset Card for "jat-dataset-tokenized"
More Information needed
jat-dataset
JAT Dataset
Dataset Description
The Jack of All Trades (JAT) dataset combines a wide range of individual datasets. It includes expert demonstrations by expert RL agents, image and caption pairs, textual data and more. The JAT dataset is part of the JAT project, which aims to build a multimodal generalist agent.
Paper: https://huggingface.co/papers/2402.09844
Usage
>>> from datasets import load_dataset
>>> dataset =… See the full description on the dataset page: https://huggingface.co/datasets/jat-project/jat-dataset.peacock-data-public-datasets-idcpaws
Dataset Card for PAWS: Paraphrase Adversaries from Word Scrambling
Dataset Summary
PAWS: Paraphrase Adversaries from Word Scrambling
This dataset contains 108,463 human-labeled and 656k noisily labeled pairs that feature the importance of modeling structure, context, and word order information for the problem of paraphrase identification. The dataset has two subsets, one based on Wikipedia and the other one based on the Quora Question Pairs (QQP) dataset.
For further… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/paws.seedance-2-prompts-datasets
🎞️ Seedance-2-prompts-datasets
🎞️ The ultimate Seedance-2 video prompt dataset (50GB+). 8100+ video generation prompts with full metadata and preview frames. Truly open source: No login, no ads, no redirection. Just pure data for AI video creators.
This project is a massive collection of prompts used for Bytedance's Seedance 2.0 and the resulting generated videos. The entire dataset exceeds 50GB and contains 8100+ videos, all structured into a comprehensive dataset.
Due… See the full description on the dataset page: https://huggingface.co/datasets/GokuScraper/seedance-2-prompts-datasets.quarel
Dataset Card for "quarel"
Dataset Summary
QuaRel is a crowdsourced dataset of 2771 multiple-choice story questions, including their logical forms.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
default
Size of downloaded dataset files: 0.63 MB
Size of the generated dataset: 1.53 MB
Total amount of disk used: 2.17 MB
An example of 'train'… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/quarel.mesh4d_datasetdataset_with_scriptThis is a test dataset.natural_questions
Dataset Card for Natural Questions
Dataset Summary
The NQ corpus contains questions from real users, and it requires QA systems to
read and comprehend an entire Wikipedia article that may or may not contain the
answer to the question. The inclusion of real user questions, and the
requirement that solutions should read an entire page to find the answer, cause
NQ to be a more realistic and challenging task than prior QA datasets.
Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/natural_questions.swe-bench-dummy-test-datasetgpt-image-2-prompts-datasets
🖼️ GPT Image 2 Prompt Dataset
🖼️ The ultimate GPT Image 2 prompt dataset (5GB+). 15,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators.
This project is a massive collection of prompts used for OpenAI's GPT Image 2 model and the resulting generated images. The entire dataset exceeds 5GB and contains 15,000+ images, all structured into a comprehensive dataset.
Due to… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/gpt-image-2-prompts-datasets.datasets-tests-compressionwikipediaWikipedia dataset containing cleaned articles of all languages.
The datasets are built from the Wikipedia dump
(https://dumps.wikimedia.org/) with one split per language. Each example
contains the content of one full Wikipedia article with cleaning to strip
markdown and unwanted sections (references, etc.).multi_dir_datasetai4g-flood-dataset
Flood Detection Dataset
Introduction
This dataset accompanies the paper Mapping global floods with 10 years of satellite radar data (Nature Communications, 2025) and contains global flood detections derived from Sentinel-1 Synthetic Aperture Radar (SAR) imagery using a deep learning change detection model. The dataset spans October 2014 – September 2024, offering a longitudinal view of flood-prone areas worldwide.
Key features:
Cloud-penetrating SAR data for consistent… See the full description on the dataset page: https://huggingface.co/datasets/ai-for-good-lab/ai4g-flood-dataset.dataset_with_data_filesEmilia-Dataset
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline.
News 🔥
2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.reddit_dataset_157
Bittensor Subnet 13 Reddit Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this dataset allows… See the full description on the dataset page: https://huggingface.co/datasets/tensorshield/reddit_dataset_157.My_Hermes_Datasettiny-supervised-datasetgrand_tour_dataset
The GrandTour Dataset
A project brought to you by RSL - ETH Zurich.
References •
Contributing •
Citation
References
Official dataset webpage: grand-tour.leggedrobotics.com
Getting started & examples: github.com/leggedrobotics/grand_tour_dataset
Boxi used to collect the data: github.com/leggedrobotics/grand_tour_box
Contributing
We warmly welcome contributions to improve and expand this project. Whether it's new examples, enhancements, or… See the full description on the dataset page: https://huggingface.co/datasets/leggedrobotics/grand_tour_dataset.dataset1synth_datasetfev_datasets
Forecast evaluation datasets
This repository contains time series datasets that can be used for evaluation of univariate & multivariate forecasting models.
The main focus of this repository is on datasets that reflect real-world forecasting scenarios, such as those involving covariates, missing values, and other practical complexities.
The datasets follow a format that is compatible with the fev package.
Data format and usage
Each dataset satisfies the following… See the full description on the dataset page: https://huggingface.co/datasets/autogluon/fev_datasets.fish_datasets_real_electrodyn_expertsys_twodim_fourierLIBERO-datasets
LIBERO Datasets
This is a repo that stores the LIBERO datasets. The structure of the dataset can be found below:
libero_object/
libero_spatial/
libero_goal/
libero_90/
libero_10/
Demonstrations of each task is stored in a hdf5 file. Please refer to download script from the official LIBERO repo for more details.
leaderboard-dataset
Arena Leaderboard Dataset
Historical snapshots of the Arena leaderboard.
Usage
from datasets import load_dataset
# Load all historical text style control data
ds = load_dataset("lmarena-ai/leaderboard-dataset", "text_style_control", split="full")
# Load the current text style control leaderboard
ds = load_dataset("lmarena-ai/leaderboard-dataset", "text_style_control", split="latest")
# Filter to overall category
ds =… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset.trackio-datasetDogSpeak_Dataset
DogSpeak: A Canine Vocalization Classification Dataset
Paper: https://dl.acm.org/doi/10.1145/3746027.3758298
Dataset: https://huggingface.co/datasets/ArlingtonCL2/DogSpeak_Dataset
Dataset Summary
DogSpeak is a large-scale canine vocalization dataset designed to advance research in animal communication and computational bioacoustics.
Unlike previous canine vocalization datasets recorded in controlled environments, DogSpeak is sourced from a large collection of… See the full description on the dataset page: https://huggingface.co/datasets/ArlingtonCL2/DogSpeak_Dataset.
