datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SPEED-Bench
📒 Blog |
📄 Paper |
🤗 Data |
⚙️ Measurement Framework
SPEED-Bench (SPEculative Evaluation Dataset) is a unified benchmark designed to evaluate speculative decoding (SD) across diverse semantic domains and realistic serving regimes, using production-grade inference engines.
It measures both acceptance-rate characteristics and end-to-end throughput, enabling fair, reproducible, and robust comparisons between SD strategies.
SPEED-Bench introduces a benchmarking ecosystem for… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/SPEED-Bench.linkwarp-speed
WarpSpeed Research Dataset
Dataset Summary
The WarpSpeed Research Dataset is a comprehensive collection of scientific research papers, experimental data, and theoretical materials focused on advanced propulsion concepts and physics principles inspired by Star Trek technologies. This dataset combines real-world physics research with theoretical frameworks to explore the possibilities of faster-than-light travel and advanced energy systems.
Data Collection and… See the full description on the dataset page: https://huggingface.co/datasets/GotThatData/warp-speed.codetorch-archiveBench2Drive-Speed
Bench2Drive-Speed
Project Page | Paper | GitHub
Bench2Drive-Speed is a closed-loop benchmark for desired-speed conditioned autonomous driving, enabling explicit control over vehicle behavior through target speed and overtake/follow commands.
The CustomizedSpeedDataset contains 2,100 CARLA driving scenarios with expert demonstrations and annotated overtake/follow commands. The released dataset includes expert target-speed signals only.
Virtual Target Speed Annotation
You… See the full description on the dataset page: https://huggingface.co/datasets/rethinklab/Bench2Drive-Speed.bus_speeds_inner_hcmlaions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuningLAION's Got Talent: Generated Voice Acting Dataset
Overview
"LAION's Got Talent" is a synthetic voice acting dataset designed to offer a broad range of emotional expressions, vocal bursts, and multi-language utterances. This dataset is a component of the BUD-E project, led by LAION with support from Intel, and aims to drive forward research in context-aware and empathetic AI voice assistants.
Updated Composition
Voices and Languages
English: 11 OpenAI voices, each… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuning.trivia_qa_tiny
Dataset Card for Dataset Name
Dataset Summary
This dataset contains 100 samples from trivia_qa dataset. It is used mainly for testing purposes.
Languages
English.
Dataset Structure
Data Instances
Total data size: 8Kb.
Data Fields
question: string feature, containing question to be answered.
`answer: string feature, answer to the question.
Data Splits
Only test split, that contains 100 rows, is supported.
cua-speedrun-trajectories
CUA Speedrun results
Trajectory archives and measured results for the benchmark sets used by CUA Speedrun.
Unanimous-295
Results on the 295-task OSWorld set used by CUA Speedrun. Except for the explicitly reported GPT-5.6 Luna aggregate row, scores are the mean normalized verifier score across the 295 tasks and may include partial credit. Times are measured task time per task.
Model
Average score
Average task time
Average cost per task
Kimi K3
85.07%… See the full description on the dataset page: https://huggingface.co/datasets/anonymousmypcbench/cua-speedrun-trajectories.llm_speedrun
LLM Speedrun token streams
Pre-tokenized training artifacts for the LLM speedrun exercises.
File
Description
Tokens
tokenizer_50M.bpe
JSON-serialized BPE tokenizer
—
fineweb-edu-10BT.shuffle.bin
Shuffled FineWeb-Edu sample/10BT token stream
9,440,023,113
smoltalk.shuffle.bin
Shuffled SmolTalk data/all token stream
875,269,408
The .bin files are headerless, little-endian unsigned 16-bit token IDs and can be memory-mapped with NumPy:
from huggingface_hub import… See the full description on the dataset page: https://huggingface.co/datasets/zkolter/llm_speedrun.NYC_traffic_speed
Data source:
NYC_Traffic_Speed
Dataset Structure
The dataset is organized into the following structure:
|-- subdataset1
| |-- raw_data # Original data files
| |-- time_series # Rule-based Imputed data files
| | |-- id_1.parquet # Time series data for each subject can be multivariate, can be in csv, parquet, etc.
| | |-- id_2.parquet
| | |-- ...
| | |-- id_info.json # Metadata for each subject
| |-- weather
| | |-- location_1
| | | |--… See the full description on the dataset page: https://huggingface.co/datasets/fidel-ts/NYC_traffic_speed.gr00t-g1-grab-bottle-right-hand-speedup-3mm-v1
Grab-Bottle (right hand) - DP-resampled, 3 mm/frame wrist target
LeRobot v2.1 dataset for the Unitree G1 right-hand bottle-grab task. Merges two
teleoperation datasets and applies Dynamic Programming (DP) resampling on the
right-wrist XYZ trajectory to normalise the action density across all episodes. It is
the training data for the
v7 GR00T N1.7 fine-tune
a speedup-resampling reference baseline (trained 2026-06-25, checkpoint-20000
only).
The key idea: instead of keeping… See the full description on the dataset page: https://huggingface.co/datasets/cloudwalk-research/gr00t-g1-grab-bottle-right-hand-speedup-3mm-v1.JDocQAThis unofficial dataset consists of QA pairs with images converted from the PDF files of JDocQA, a dataset focusing on chart and table understanding.
The conversion was performed using pdf2image.
The original dataset includes 1,176 examples, but 12 examples could not be converted into images. As a result, this image dataset consists of 1,164 examples in total.
We are uploading it here for use in the evaluation of llm-jp-eval-mm.
Please see the official github repo… See the full description on the dataset page: https://huggingface.co/datasets/speed/JDocQA.so100_grasp_pink_blockThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 43255,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/speedyyoshi/so100_grasp_pink_block.gr00t-g1-grab-bottle-right-hand-speedup-3mm-cycle-removed-v1
Grab-Bottle (right hand) - DP-resampled 3 mm/frame + cycle removal
LeRobot v2.1 dataset for the Unitree G1 right-hand bottle-grab task. Extends the
DP-resampled variant
by automatically detecting and removing return cycles in wrist-space - segments where
the arm travels away from a point and then returns to (approximately) the same position.
Such cycles teach the policy to backtrack rather than commit to a goal-directed motion.
No fine-tune has been produced from this variant… See the full description on the dataset page: https://huggingface.co/datasets/cloudwalk-research/gr00t-g1-grab-bottle-right-hand-speedup-3mm-cycle-removed-v1.piper-speedup-evals
PiPER speed-up evals
Real-robot policy rollouts on the PiPER bimanual cell (plate pick, hand-over, place),
run from the operator stack's Policy panel at execution speed multipliers of 1x to 6x.
One folder per policy, one sub-folder per run (start time, IST): the three camera clips
(top, left-arm, right-arm; re-encoded from the recorded chunks with every frame at its
capture time, so they play in real time; clip t = 0 is 1 s before the run start and
frame_ts.json lists each… See the full description on the dataset page: https://huggingface.co/datasets/PranayTest/piper-speedup-evals.WAON
WAON: Large-Scale and High-Quality Japanese Image-Text Pair Dataset for Vision-Language Models
|
🤗 HuggingFace
|
📄 Paper
|
🧑💻 Code
|
Introduction
WAON is a Japanese (image, text) pair dataset containing approximately 155M examples, crawled from Common Crawl.
It is built from snapshots taken in 2025-18, 2025-08, 2024-51, 2024-42, 2024-33, and 2024-26.
The dataset is high-quality and diverse, constructed through a sophisticated… See the full description on the dataset page: https://huggingface.co/datasets/speed/WAON.NYC_traffic_speed
WIATS: Weather-centric Intervention-Aware Time Series Multimodal Dataset
Data source:
NYC_Traffic_Speed
Dataset Structure
The dataset is organized into the following structure:
|-- subdataset1
| |-- raw_data # Original data files
| |-- time_series # Rule-based Imputed data files
| | |-- id_1.parquet # Time series data for each subject can be multivariate, can be in csv, parquet, etc.
| | |-- id_2.parquet
| | |-- ...
| | |-- id_info.json… See the full description on the dataset page: https://huggingface.co/datasets/VEWOXIC/NYC_traffic_speed.gigaword_tiny
Dataset Card for Dataset Name
This is a tiny version of https://huggingface.co/datasets/Harvard/gigaword, used for testing purposes.
Dataset Details
Dataset Description
This is a tiny version of https://huggingface.co/datasets/Harvard/gigaword.
It was created by selecting only first 100 samples from each split.
Language(s) (NLP): English
Dataset Sources [optional]
Uses
It is supposed be used only for testing purposes.… See the full description on the dataset page: https://huggingface.co/datasets/SpeedOfMagic/gigaword_tiny.Speedhunters_cleanthe missing articles.It is already not in the website before I did the backups.here is a russian translated version
https://carakoom.com/blog/10328
VR_H31_bodyshop_place_part2_pose5_stereo_rvt_trym_speedup120260428_VR_H31_bodyshop_place_part2_pose46_nav_stereo_rvt_speedup1speedrunbench
SpeedrunBench
LLM agents optimizing speedruns across multiple games. Each run ships as a silent video of the run that
was scored, and (except for Tuxemon) the input tape that produced it: the frame to goal and lower
is always better.
config
game
metric
supertux
SuperTux
frames to goal
astray
Astray
frames to maze exit
tuxemon
Tuxemon
frames to first gym
pokemon_blue
Pokémon Blue
frames to first badge
sml
Super Mario Land
frames to clear 1-1
smb
Super Mario… See the full description on the dataset page: https://huggingface.co/datasets/PatronusAI/speedrunbench.Y_speed_time20260326_VR_H31_bodyshop_pick_part2_nav_stereo_rvt_speedup1ontonotes_english
Dataset Card for ontonotes_english
Dataset Summary
This is preprocessed version of what I assume is OntoNotes v5.0.
Instead of having sentences stored in files, files are unpacked and sentences are the rows now. Also, fields were renamed in order to match conll2003.
The source of data is from private repository, which in turn got data from another public repository, location of which is unknown :)
Since data from all repositories had no license (creator of the private… See the full description on the dataset page: https://huggingface.co/datasets/SpeedOfMagic/ontonotes_english.20260413_VR_H31_bodyshop_place_part2_pose25_nav_stereo_rvt_speedup1vrh3_placing_part_speedup1_virtual_100peval1_task1_speed_coloredThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/robot-learning-group47/eval1_task1_speed_colored.gr00t-g1-grab-bottle-right-hand-speedup-2mm-v3
Grab-Bottle (right hand) - DP-resampled, 2 mm/frame wrist target (on curated v2)
LeRobot v2.1 dataset for the Unitree G1 right-hand bottle-grab task. Applies
Dynamic Programming (DP) resampling at a finer 2 mm/frame wrist target on top
of the pre-curated gr00t-g1-grab-bottle-right-hand-v2
(wandering-removed) dataset. It is the training data for the
v9 GR00T N1.7 fine-tune
(trained 2026-06-26, checkpoint-10000 + checkpoint-20000).
The key idea: instead of keeping every raw frame… See the full description on the dataset page: https://huggingface.co/datasets/cloudwalk-research/gr00t-g1-grab-bottle-right-hand-speedup-2mm-v3.
