datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fev_datasets
Forecast evaluation datasets
This repository contains time series datasets that can be used for evaluation of univariate & multivariate forecasting models.
The main focus of this repository is on datasets that reflect real-world forecasting scenarios, such as those involving covariates, missing values, and other practical complexities.
The datasets follow a format that is compatible with the fev package.
Data format and usage
Each dataset satisfies the following… See the full description on the dataset page: https://huggingface.co/datasets/autogluon/fev_datasets.leaderboard-dataset
Arena Leaderboard Dataset
Historical snapshots of the Arena leaderboard.
Usage
from datasets import load_dataset
# Load all historical text style control data
ds = load_dataset("lmarena-ai/leaderboard-dataset", "text_style_control", split="full")
# Load the current text style control leaderboard
ds = load_dataset("lmarena-ai/leaderboard-dataset", "text_style_control", split="latest")
# Filter to overall category
ds =… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset.chronos_datasets
Chronos datasets
Time series datasets used for training and evaluation of the Chronos forecasting models.
Note that some Chronos datasets (ETTh, ETTm, brazilian_cities_temperature and spanish_energy_and_weather) that rely on a custom builder script are available in the companion repo autogluon/chronos_datasets_extra.
See the paper for more information.
Data format and usage
The recommended way to use these datasets is via https://github.com/autogluon/fev.
All datasets… See the full description on the dataset page: https://huggingface.co/datasets/autogluon/chronos_datasets.trex_dataset
T-Rex Dataset
A large-scale, tactile-reactive bimanual manipulation dataset, collected via teleoperation on a
Dexmate Vega-1 robot with two Sharpa Wave dexterous hands. Stored as a
LeRobotDataset v3.0.
🌐 Project Page · ✍️ Paper (arXiv) · 💻 Code (T-Rex) · 🚀 Dataset Quickstart · 📓 Colab notebook
One episode from each of 20 motor primitives (head-camera view, cropped to the workspace), each with a different object.
Teleoperation setup: Manus gloves + VIVE… See the full description on the dataset page: https://huggingface.co/datasets/zekaiwang/trex_dataset.CADS-dataset
CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography
Overview
CADS is a robust, fully automated framework for segmenting 167 anatomical structures in Computed Tomography (CT), spanning from head to knee regions across diverse anatomical systems.
The framework consists of two main components:
CADS-dataset:
22,022 CT volumes with complete annotations for 167 anatomical structures.
Most extensive whole-body CT dataset… See the full description on the dataset page: https://huggingface.co/datasets/huggingface/CADS-dataset.MolmoAct2-BimanualYAM-DatasetThis dataset was created using LeRobot.
MolmoAct2-BimanualYAM Dataset
This repository is the merged ckpt / merged LeRobot dataset artifact for the MolmoAct2-BimanualYAM Dataset, a large-scale collection of bimanual robot manipulation demonstrations collected for MolmoAct2. Across the full collection, MolmoAct2-BimanualYAM contains more than 720 hours of training demonstrations spanning diverse tabletop manipulation tasks.
Language Annotations
This dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoAct2-BimanualYAM-Dataset.show3d-dataset
SHOW3D: Capturing Scenes of 3D Hands and Objects in the Wild
Patrick Rim, Kevin Harris, Braden Copple, Shangchen Han, Xu Xie, Ivan Shugurov, Sizhe An, He Wen,
Alex Wong, Tomas Hodan, and Kun He
CVPR 2026; https://arxiv.org/abs/2603.28760
News
September 18, 2026: Released synchronized exocentric views and camera calibrations.
SHOW3D is a large-scale multi-view dataset of hand–object interactions captured in the wild.
It is intended to advance research on… See the full description on the dataset page: https://huggingface.co/datasets/facebook/show3d-dataset.go_emotions
Dataset Card for GoEmotions
Dataset Summary
The GoEmotions dataset contains 58k carefully curated Reddit comments labeled for 27 emotion categories or Neutral.
The raw data is included as well as the smaller, simplified version of the dataset with predefined train/val/test
splits.
Supported Tasks and Leaderboards
This dataset is intended for multi-class, multi-label emotion classification.
Languages
The data is in English.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/go_emotions.short_video_ocr_dataset
Short Video OCR / ASR Dataset
An actively curated research dataset for building OCR, ASR, subtitle-alignment,
and video-transcript pipelines for short social videos. It combines source
videos and extracted frames with human review artifacts and model-generated
text candidates. The primary languages are Ukrainian and Russian; English or
mixed-language content may also occur.
Status: work in progress. Model outputs and pseudo-label candidates are
not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.car-bench-dataset
CAR-Bench Dataset
CAR-Bench is a benchmark for evaluating AI voice assistants in a realistic automotive (car) environment.
It tests an agent's ability to correctly use vehicle control tools, handle disambiguation, and avoid hallucinations.
Dataset Structure
The dataset is organized into task configs and mock data configs:
Tasks
Each task defines a user persona, an instruction, the initial vehicle/environment context, and the ground-truth sequence of tool-call… See the full description on the dataset page: https://huggingface.co/datasets/johanneskirmayr/car-bench-dataset.dataset-2024-12-18dataset-2024-12-17CADS-dataset
CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography
Overview
CADS is a robust, fully automated framework for segmenting 167 anatomical structures in Computed Tomography (CT), spanning from head to knee regions across diverse anatomical systems.
The framework consists of two main components:
CADS-dataset:
22,022 CT volumes with complete annotations for 167 anatomical structures.
Most extensive whole-body CT dataset… See the full description on the dataset page: https://huggingface.co/datasets/sunghong/CADS-dataset.fmb_dataset_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 8612,
"total_frames": 1137459,
"total_tasks": 24,
"total_videos": 34448,
"total_chunks": 9,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:8612"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/fmb_dataset_lerobot.dataset-2024-12-16minuszero-indian-autonomous-driving-dataset-v2
INDUS-AD: Indian Dataset of Unstructured Urban Scenes for Autonomous Driving
Overview
INDUS-AD is the largest publicly released Indian autonomous-driving dataset for end-to-end autonomous-driving research. Its name expands to Indian Dataset of Unstructured Urban Scenes for Autonomous Driving.
This gated dataset is the decoded companion to the Minus Zero Indian Urban Autonomous Driving Dataset. It provides directly usable camera MP4s, normalized sensor tables… See the full description on the dataset page: https://huggingface.co/datasets/gagandeepreehal/minuszero-indian-autonomous-driving-dataset-v2.sports-trends-dataset
⚽🏀🎾🏏 Sports-Trends Dataset
A leakage-safe, multi-sport match data lake — raw fixtures → engineered features → training splits.
The data backbone of Ruslan Magana Sports Intelligence — refreshed automatically every day.
TL;DR — A continuously-updated, medallion-architecture data lake for football,
basketball, tennis and cricket: immutable raw ingests, cleaned/standardized layers, an
engineered feature store, and ready-to-train chronological splits in… See the full description on the dataset page: https://huggingface.co/datasets/ruslanmv/sports-trends-dataset.spotify-tracks-dataset
Content
This is a dataset of Spotify tracks over a range of 125 different genres. Each track has some audio features associated with it. The data is in CSV format which is tabular and can be loaded quickly.
Usage
The dataset can be used for:
Building a Recommendation System based on some user input or preference
Classification purposes based on audio features and available genres
Any other application that you can think of. Feel free to discuss!
Column… See the full description on the dataset page: https://huggingface.co/datasets/maharshipandya/spotify-tracks-dataset.music-arena-dataset
Music Arena Dataset
This is the official dataset from Music Arena, an open platform for evaluating text-to-music (TTM) models.
How to Download (Recommended Method)
The most reliable way to get a complete local copy of all files, including the entire audio collection, is to clone the repository directly using Git. This method is ideal for offline access and workflows that require direct file manipulation.
Note: This repository uses Git LFS (Large File Storage) to… See the full description on the dataset page: https://huggingface.co/datasets/music-arena/music-arena-dataset.vifailback-dataset-lerobot
ViFailback Dataset — LeRobot
This repository is a LeRobot v2.1 conversion of the trajectory portion of sii-rhos-ai/ViFailback-Dataset, introduced in the CVPR 2026 paper Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbols.
ViFailback contains real-world ALOHA dual-arm manipulation trajectories designed for studying failure diagnosis, failure localization, corrective guidance, recovery, and learning… See the full description on the dataset page: https://huggingface.co/datasets/sii-rhos-ai/vifailback-dataset-lerobot.opencs2_dataset_demo
HLTV CS2 Demos Dataset
Counter-Strike 2 match demos scraped from HLTV.org
plus a compact per-map analysis JSON. Each row of the metadata Parquet is
one .dem file (one played CS2 map); a best-of-3 match contributes 2
or 3 rows depending on whether it went 2-0 or 2-1.
Parquet holds everything you typically filter on: map_name,
patch_version, rounds_played, per-player kast / adr / rating,
every kill tick with weapon + headshot, every round's winner + end reason.
.dem binaries live… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/opencs2_dataset_demo.Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/SHSLab/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.furniture_bench_dataset_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "franka",
"total_episodes": 5100,
"total_frames": 3948057,
"total_tasks": 9,
"total_videos": 10200,
"total_chunks": 6,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:5100"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/furniture_bench_dataset_lerobot.transformers-pr-slop-dataset
Transformers PR Slop Dataset
Normalized snapshots of issues, pull requests, comments, reviews, and linkage data from huggingface/transformers.
Files:
issues.parquet
pull_requests.parquet
comments.parquet
issue_comments.parquet (derived view of issue discussion comments)
pr_comments.parquet (derived view of pull request discussion comments)
pr_files.parquet
pr_diffs.parquet
reviews.parquet
review_comments.parquet
links.parquet
events.parquet
Use:
duplicate PR and issue analysis… See the full description on the dataset page: https://huggingface.co/datasets/burtenshaw/transformers-pr-slop-dataset.soccer-dataset
Global Football (Soccer) Data Lake
Cleaned, deduplicated, quality-gated football match data for BTTS / goals modelling.
Sources: API-Football + football-data.co.uk. Pipeline & docs:
https://github.com/eatpizzanot/soccer-dataset
673,966 fixtures (644,901 played), 271 leagues,
11,104 teams, 2008-06-07 - 2027-06-06.
BTTS base rate 0.5063. xG fake-zeros removed; known_at leakage guard;
12-dimension QA gate (QUALITY_REPORT.md).
Caveats: league history is uneven — check… See the full description on the dataset page: https://huggingface.co/datasets/eatpizzanot/soccer-dataset.opencs2_dataset
OpenCS2 - POV Renders
Browse with the OpenCS2 Viewer - every match, map and round, with all 10 player POVs synced on one timeline.
Tick-aligned Counter-Strike 2 POV training clips, rendered from
blanchon/cs2_dataset_demo. Each row
in the main table is one player's perspective for one round; ten POVs per round share the same tick
clock.
Per POV round:
Video - 1280x720 @ 32 fps, near-lossless H.264, faststart, muxed with audio.
Audio - per-player stereo, mixed from that… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/opencs2_dataset.dataset-the-stack-v2-dedup-sub
The Stack v2 Subset with File Contents (Python, Java, JavaScript, C, C++)
TempestTeam/dataset-the-stack-v2-dedup-sub
Dataset Summary
This dataset is a language-filtered and self-contained subset of bigcode/the-stack-v2-dedup, part
of the BigCode Project.
It contains only files written in the following programming languages:
Python 🐍
Java ☕
JavaScript 📜
C ⚙️
C++ ⚙️
Unlike the original dataset, which only includes metadata and Software Heritage IDs, this subset includes… See the full description on the dataset page: https://huggingface.co/datasets/Cyrile/dataset-the-stack-v2-dedup-sub.fmars-dataset
FMARS: Foundation Model Annotations for Remote Sensing Images
FMARS is a large-scale dataset of Very High Resolution (VHR) remote sensing images with annotations generated using Vision Foundation Models.
The dataset focuses on disaster management applications and provides pre-event imagery and annotations for major crisis events worldwide from 2021 to 2023.
Paper: https://arxiv.org/abs/2405.20109
Dataset Features
VHR Imagery: The dataset uses pre-event VHR… See the full description on the dataset page: https://huggingface.co/datasets/links-ads/fmars-dataset.marl-gpt-datasets
MARL-GPT Datasets
Offline expert trajectories from “MARL-GPT: Foundation Model for Multi-Agent Reinforcement Learning”.
Environments
This dataset includes trajectories from the three evaluation domains used in MARL-GPT: SMACv2 (StarCraft multi-agent combat), Google Research Football (GRF), and POGEMA (partially observable multi-agent pathfinding on grids).
Format
Trajectories are stored sequentially (no shuffling). Use the done flag to split the stream into… See the full description on the dataset page: https://huggingface.co/datasets/nortem/marl-gpt-datasets.Peregrine-Dataset-v2023-11
Peregrine v2023-11 — Layer-wise L-PBF Imaging Dataset
A HuggingFace-formatted conversion of the Peregrine v2023-11 dataset released by Oak Ridge National Laboratory's (ORNL) Manufacturing Demonstration Facility (MDF). The dataset contains layer-wise in-situ imaging, anomaly segmentation masks, laser scan paths, and process sensor data from 5 Laser Powder Bed Fusion (L-PBF) builds of stainless steel 316L.
Original dataset: Layer-wise Imaging Dataset from Powder Bed Additive… See the full description on the dataset page: https://huggingface.co/datasets/ppak10/Peregrine-Dataset-v2023-11.
