datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai2_arc
Dataset Card for "ai2_arc"
Dataset Summary
A new dataset of 7,787 genuine grade-school level, multiple-choice science questions, assembled to encourage research in
advanced question-answering. The dataset is partitioned into a Challenge Set and an Easy Set, where the former contains
only questions answered incorrectly by both a retrieval-based algorithm and a word co-occurrence algorithm. We are also
including a corpus of over 14 million science sentences… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ai2_arc.video-vec2wav2-tokenizer
video-vec2wav2-tokenizer
Production-ready pipeline (Python package video_vec2wav2_tokenizer, CLI command
video2dataset) that turns a folder of videos into clean AI training datasets
for speech recognition (ASR) and text-to-speech (TTS).
videos ──► audio (16 kHz mono PCM) ──► whisper transcript ──► clips ──► metadata.csv / dataset.jsonl / tts_metadata.csv / report.json
Video processing — recursive scan of mp4 / mkv / avi / mov / webm, FFmpeg
audio extraction to mono ·… See the full description on the dataset page: https://huggingface.co/datasets/k9cli/video-vec2wav2-tokenizer.osworld_v2_assetsps2_hf210Kh-RealOmin-OpenData
Boasting over 13,000 hours of cumulative data and 5 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry.
Update Notes:Stage 3 data upload completed.
13,000+ hours of pure dual-hand data with frame-level alignment latency < 1ms
Full high-precision trajectory reconstruction, breaking the limit of superficial open source, fully ready-to-use
3,000+ contributors and 10,000+ real household scenarios with exceptional diversity… See the full description on the dataset page: https://huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData.OpenAoE-2000h
Open-AoE — Egocentric Hand Manipulation Dataset
Release Roadmap
Tier
Duration
Status
nano
~3 h
✅ Released
tiny
~100 h
✅ Released
full
2000 h
🚧 Uploading
Release notes
2026-07-30: Removed samples flagged in PR #1 for camera-intrinsics vs. video-resolution mismatches.
2026-07-31: Uploaded ~323h of data.
2026-08-12: Uploaded ~694h of data.
2026-09-03: Uploaded ~189h of data.
Additional data for the full ~2000h release is still… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/OpenAoE-2000h.video-vec2wav2-tokenizer-2
video-vec2wav2-tokenizer-2
Version 2 - continuation shard of the video-to-AI-dataset tokenizer project.
Production-ready pipeline (Python package video_vec2wav2_tokenizer, CLI command
video2dataset) that turns a folder of videos into clean AI training datasets
for speech recognition (ASR) and text-to-speech (TTS).
videos ──► audio (16 kHz mono PCM) ──► whisper transcript ──► clips ──► metadata.csv / dataset.jsonl / tts_metadata.csv / report.json
Video processing —… See the full description on the dataset page: https://huggingface.co/datasets/k9cli/video-vec2wav2-tokenizer-2.ps2_hf1pspflatpaklatent_worker_early-a2_06latent_worker_early-a2_03challenge_data
PrimeBot Household Bimanual Manipulation Challenge Dataset
中文 | English
中文
目录
关于我们
更新日志
真机遥操作数据
训练集说明
验证集说明
数据集字段说明
URDF
图像
语言指令
本体感知与动作
机器人推理接口
UMI数据
数据概览
目录结构
数据集字段说明
图像
本体感知与动作
索引字段
标注与 IMU
关于我们
我们来自上纬新材-启元研究院,我们的使命是加速个人机器人时代到来,加速家用机器人时代到来。我们开源高质量面向家庭操作的双臂操作数据集,同时开放机器人硬件描述以供可视化、可复现研究。
如果本数据集对您的工作有帮助,感谢引用:
@misc{xu2026scalingbimanualhouseholdmanipulation,
title={Scaling Bimanual Household Manipulation from 1,500… See the full description on the dataset page: https://huggingface.co/datasets/challenge-2026/challenge_data.latent_worker_early-a2_00latent_worker_early-a2_08AgiBotWorld2026
AgiBot World 2026
Real-World Embodied Intelligence Dataset
Overview
As robotics research advances into real-world scenarios, the demand for authentic, high-quality data has become increasingly urgent. Following AGIBOT WORLD's "ImageNet moment," we now release the AGIBOT WORLD 2026 dataset. Built upon massive real-world scenes, it systematically spans pivotal research directions in embodied intelligence, designed to power the next generation of… See the full description on the dataset page: https://huggingface.co/datasets/agibot-world/AgiBotWorld2026.latent_worker_early-a2_02PIN-200M
PIN-200M
A mini version of "PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents"
Paper: https://arxiv.org/abs/2406.13923
This dataset contains around 200M samples in PIN format, with around 312 TB storage.
🚀 News
[ 2025.09.22 ] !NEW! 🔥 We have completed the final version of the PIN-200M dataset and conducted some simple statistics on it.
[ 2024.12.06 ] !NEW! 🔥 We have updated the quality signals, enabling a swift assessment of whether a sample meets… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/PIN-200M.latent_worker_early-a2_01latent_worker_early-a2_04latent_worker_early-a2_07arxiv-cs-2020-2025-pdfsvideo-vec2wav2-tokenizer-3
video-vec2wav2-tokenizer-3
Version 3 - continuation shard of the video-to-AI-dataset tokenizer project.
Version 2 - continuation shard of the video-to-AI-dataset tokenizer project.
Production-ready pipeline (Python package video_vec2wav2_tokenizer, CLI command
video2dataset) that turns a folder of videos into clean AI training datasets
for speech recognition (ASR) and text-to-speech (TTS).
videos ──► audio (16 kHz mono PCM) ──► whisper transcript ──► clips ──►… See the full description on the dataset page: https://huggingface.co/datasets/k9cli/video-vec2wav2-tokenizer-3.dcvlm-baseline-200b
DCVLM-Baseline (200B tokens)
DCVLM-Baseline is the reference training mixture from our DataComp-VLM paper.
It is a pre-mixed, decontaminated, ready-to-train multimodal pretraining dataset, materialized as flat
WebDataset tar shards so it can be consumed by any training
stack.
This is a 200B-token dataset release consisting of 103,985,276 samples, curated from our DCVLM-large data pool.
A smaller 6.25B-token version is also available.
⚠️ NOTE: The training data is the WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm-baseline-200b.dom-pi-pdfs-2025
DOM-PI 2025 — PDFs-fonte (Diário Oficial dos Municípios do Piauí)
PDFs originais das publicações de 2025 do Diário Oficial dos Municípios do Piauí,
organizados por Território de Desenvolvimento. São a fonte da qual o corpus textual foi
extraído por OCR/parsing. 41.617 PDFs · ~70 GB.
Dataset de texto derivado (carregável, com limpeza e tiers de qualidade):
gutoportelaa/dom-pi-corpus-2025.
Cobertura de PDFs (parcial): presentes 7 territórios — tabuleiros_alto_parnaiba… See the full description on the dataset page: https://huggingface.co/datasets/gutoportelaa/dom-pi-pdfs-2025.seedance-2-prompts-datasets
🎞️ Seedance-2-prompts-datasets
🎞️ The ultimate Seedance-2 video prompt dataset (50GB+). 8100+ video generation prompts with full metadata and preview frames. Truly open source: No login, no ads, no redirection. Just pure data for AI video creators.
This project is a massive collection of prompts used for Bytedance's Seedance 2.0 and the resulting generated videos. The entire dataset exceeds 50GB and contains 8100+ videos, all structured into a comprehensive dataset.
Due… See the full description on the dataset page: https://huggingface.co/datasets/GokuScraper/seedance-2-prompts-datasets.HPLT2.0_cleanedNB: HPLT2.0 is now superseded by a newer release:
HPLT3.0
We recommed switching to v3.0, unless you have a compelling reason to stay on 2.0.
This is a large-scale collection of web-crawled documents in 191 world languages, produced by the HPLT project.
The source of the data is mostly Internet Archive with some additions from Common Crawl.
For a detailed description of the dataset, please refer to our website and our pre-print.
The Cleaned variant of HPLT Datasets v2.0
This is… See the full description on the dataset page: https://huggingface.co/datasets/HPLT/HPLT2.0_cleaned.HelpSteer2
HelpSteer2: Open-source dataset for training top-performing reward models
HelpSteer2 is an open-source Helpfulness Dataset (CC-BY-4.0) that supports aligning models to become more helpful, factually correct and coherent, while being adjustable in terms of the complexity and verbosity of its responses.
This dataset has been created in partnership with Scale AI.
When used to tune a Llama 3.1 70B Instruct Model, we achieve 94.1% on RewardBench, which makes it the best Reward Model as… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/HelpSteer2.latent_worker_early3_2fractal20220817_data_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "google_robot",
"total_episodes": 87212,
"total_frames": 3786400,
"total_tasks": 599,
"total_videos": 87212,
"total_chunks": 88,
"chunks_size": 1000,
"fps": 3,
"splits": {
"train": "0:87212"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/fractal20220817_data_lerobot.
