datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aloha_sim_transfer_cube_humanThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 20000,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_sim_transfer_cube_human.aloha_sim_insertion_humanThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 25000,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_sim_insertion_human.esg_reports_human_labeled_v2
Vidore Benchmark 2 - ESG Human Labeled
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of ESG reports from the fast food industry.
Dataset Summary
Each query is in english.
This dataset provides a focused benchmark for visual retrieval tasks related to ESG reports for the fast food industry. It includes a curated set of documents, queries, relevance judgments (qrels), and page… See the full description on the dataset page: https://huggingface.co/datasets/vidore/esg_reports_human_labeled_v2.humanoid-robots-training-dataset
Dynamic Intelligence — Humanoid Robot Training Dataset
A first-person (egocentric) video dataset of human hand manipulation, designed for training humanoid robot policies via imitation learning. Each episode captures a person performing an everyday household task — folding clothes, moving dishes, opening doors — filmed from a head-mounted iPhone using its built-in LiDAR and depth sensors.
The dataset pairs each video with frame-level 3D hand tracking and camera pose data, giving… See the full description on the dataset page: https://huggingface.co/datasets/DynamicIntelligence/humanoid-robots-training-dataset.arena-human-preference-55kDataset for Kaggle competition on predicting human preference on Chatbot Arena battles.
The training dataset includes over 55,000 real-world user and LLM conversations and user preferences across over 70 state-of-the-art LLMs, such as GPT-4, Claude 2, Llama 2, Gemini, and Mistral models.
Each sample represents a battle consisting of 2 LLMs which answer the same question, with a user label of either prefer model A, prefer model B, tie, or tie (both bad).
Citation
Please cite the… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/arena-human-preference-55k.mt_bench_human_judgments
Content
This dataset contains 3.3K expert-level pairwise human preferences for model responses generated by 6 models in response to 80 MT-bench questions.
The 6 models are GPT-4, GPT-3.5, Claud-v1, Vicuna-13B, Alpaca-13B, and LLaMA-13B. The annotators are mostly graduate students with expertise in the topic areas of each of the questions. The details of data collection can be found in our paper.
Agreement Calculation
This Colab notebook shows how to compute the… See the full description on the dataset page: https://huggingface.co/datasets/lmsys/mt_bench_human_judgments.HumanTracker
Dataset Card for HumanTracker
Project page · Paper · Code
HumanTracker is a humanoid motion-tracking benchmark. This release contains two complementary subsets:
motions/ — the evaluation test split: retargeted 29-DoF reference trajectories, grouped into four motion families.
preference_pair/ — 6,000 human preference pairs, each stored with the two tracker rollouts that were compared and the source-motion clip they track.
The evaluation harness and HumanScore reward model live… See the full description on the dataset page: https://huggingface.co/datasets/GalaxyGeneralRobotics/HumanTracker.Humanoid-Everyday-G1robocasa_target_human_unifiedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "robocasa",
"total_episodes": 25307,
"total_frames": 14957899,
"total_tasks": 50,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:25307"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/robocasa_target_human_unified.Humanoid-Everyday-G1arena-human-preference-100k
Overview
This dataset contains leaderboard conversation data collected between June 2024 and August 2024.
It includes English human preference evaluations used to develop Arena Explorer.
Additionally, we provide an embedding file, which contains precomputed embeddings for the English conversations.
These embeddings are used in the topic modeling pipeline to categorize and analyze these conversations.
For a detailed exploration of the dataset and analysis methods, refer to the… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/arena-human-preference-100k.text-2-video-human-preferences
Rapidata Video Generation Preference Dataset
This dataset was collected in ~12 hours using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
The data collected in this dataset informs our text-2-video model benchmark. We just started so currently only two models are represented in this set:
Sora
Hunyouan
Pika 2.0
Runway ML Alpha
Luma Ray 2
Explore our latest model rankings on our website.
If you get value from this dataset and would… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences.text-2-video-human-preferences-wan2.1
Rapidata Video Generation Alibaba Wan2.1 Human Preference
If you get value from this dataset and would like to see more in the future, please consider liking it.
This dataset was collected in ~1 hour total using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Overview
In this dataset, ~45'000 human annotations were collected to evaluate Alibaba Wan 2.1 video generation model on our benchmark. The up to date benchmark… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-wan2.1.PPE-Human-Preference-V1
Overview
This contains the human preference evaluation set for Preference Proxy Evaluations.
This dataset is meant for benchmarking and evaluation, not for training.
Paper
Code
License
User prompts are licensed under CC-BY-4.0, and model outputs are governed by the terms of use set by the respective model providers.
Citation
@misc{frick2024evaluaterewardmodelsrlhf,
title={How to Evaluate Reward Models for RLHF},
author={Evan Frick and Tianle Li and… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/PPE-Human-Preference-V1.robocasa_pretrain_human300_v4_annotated5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.images.robot0_agentview_left": {
"dtype": "video",
"shape": [
256,
256,
3
],
"names": [
"height",
"width",
"channel"
],
"video_info": {… See the full description on the dataset page: https://huggingface.co/datasets/pepijn223/robocasa_pretrain_human300_v4_annotated5.gspc-human-labour-index
GSPC — labour components facts (Eurostat)
SWIFT census (live): https://councilof.ai/api/swift
XRPL reader (live): https://councilof.ai/api/xrpl
Live axis name: labour-components — MEASURED as two labour series (deterministic-facts, n=2). Not an index. No composite. C-2026-0826-05: do not restore MEASURED-INDEX-v0.1.
Legacy Hub slug kept for inbound links. Cite the live axis name. Empty cells that the GET does not fill stay empty.
Council of AI · CSOAI Ltd (GB, Companies House… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-human-labour-index.gspc-humanoid-labour-index
GSPC — humanoid labour index facts (Disclosure)
SWIFT census (live): https://councilof.ai/api/swift
XRPL reader (live): https://councilof.ai/api/xrpl
MEASURED financial/domain axis as disclosure facts on 8 frozen URLs (n=8). Not a model leaderboard. No accuracy, no fleet, no leader. Cells the live GET leaves empty stay empty — never invent hours/incidents.
Live status is the humanoid-labour-index row on GET https://councilof.ai/api/gspc. Not a certificate.
Council of AI ·… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-humanoid-labour-index.human_genomeQIT
QIT Humanize-Physic Formalizations and Proofs
QIT (Quantum Information Theory) is a blind benchmark for formalizing theorems in quantum information. It evaluates whether an AI agent can faithfully translate natural-language and TeX problem statements into Lean 4 theorems and then construct formal proofs checked by the Lean kernel. Its 40 tasks cover quantum channels and Choi representations, entropy and coding, mixed-unitary obstructions and symmetry, norm and fidelity tools… See the full description on the dataset page: https://huggingface.co/datasets/humanfia-lab/QIT.robocasa_pretrain_human300_v4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.images.robot0_agentview_left": {
"dtype": "video",
"shape": [
256,
256,
3
],
"names": [
"height",
"width",
"channel"
],
"video_info": {… See the full description on the dataset page: https://huggingface.co/datasets/pepijn223/robocasa_pretrain_human300_v4.humans-top
humans.top — LIVE Global ranking of influential people (open dataset)
This dataset ranks real, named living people by global influence — e.g. #1
Donald Trump, #2 Xi Jinping, #3 Vladimir Putin, alongside figures like Elon Musk,
Narendra Modi and Lionel Messi. Every row is a person: their live influence
rank, a concise biography in 15 languages, and Wikidata / Wikipedia links.
Published from the website humans.top (.top is the
domain name).
Available on (identical CC0… See the full description on the dataset page: https://huggingface.co/datasets/dsfox/humans-top.human_ref_dna
Dataset Card for "human_ref_dna"
More Information needed
text-2-video-human-preferences-seedance-1-pro
Rapidata Video Generation Seedance 1 Pro Human Preference
In this dataset, ~60k human responses from ~20k human annotators were collected to evaluate Seedance 1 Pro video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-seedance-1-pro.egopi_latal_humanhumanoid-everyday-stepit
Humanoid-Everyday · stepit action(G1 子集)
一个自包含的标准 LeRobot v2.1 数据集,由原始 Humanoid-Everyday 数据集的 Unitree G1 子集
重新表达而来。与原始数据唯一的区别是 action 列被替换成了一个 41 维、包含完整 base 状态
(位置 + 姿态 + 线速度)的动作向量;其余所有列均逐字节保持不变,可直接用标准 LeRobot 加载器读取。
动作空间:41 维 = base(10) + G1 身体 29 关节 + 双手开合 2;抽取相应切片即可喂给
stepit(Unitree G1 29-DOF)控制器(见下文 stepit qpos)。
规模:4068 episodes / 1,781,092 frames / 246 tasks / 9 chunks,fps = 30,robot_type = g1。
体积:≈ 394 GB(parquet ≈ 391 GB + 视频 ≈ 3 GB + meta ≈ 2 MB)。
原始数据是 G1/H1 混合的(共… See the full description on the dataset page: https://huggingface.co/datasets/UsanoCoCr/humanoid-everyday-stepit.text-to-speech-human-preferences-315k
Text-to-speech human preferences: 315K votes across 15 models
This gated dataset contains the evaluation record behind Datapoint Audio
Bench: 315,000 eligible pairwise votes comparing 15 text-to-speech
models in a complete round-robin over 300 English prompts. The prompt set
covers eight practical voice-agent categories, and every generated sample is
included as a typed audio record.
The source evaluation collected 357,651 completed responses. The published
benchmark excluded… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-to-speech-human-preferences-315k.PLM-Video-Human
Dataset Card for PLM-Video Human
PLM-Video-Human is a collection of human-annotated resources for training Vision Language Models,
focused on detailed video understanding. Training tasks include: fine-grained open-ended question answering (FGQA), Region-based Video Captioning (RCap),
Region-based Dense Video Captioning (RDCap) and Region-based Temporal Localization (RTLoc).
[📃 Tech Report]
[📂 Github]
Dataset Structure
Fine-Grained Question Answering (FGQA)… See the full description on the dataset page: https://huggingface.co/datasets/facebook/PLM-Video-Human.text-2-video-human-preferences-veo3
Rapidata Video Generation Veo 3 Human Preference
In this dataset, ~46k human responses from ~20k human annotators were collected to evaluate Veo3 video generation model on our benchmark. This dataset was collected in roughly 35 minutes using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-veo3.humanego_serve_bread_lingbot_lerobot_with_latents
HumanEgo Serve Bread LingBot LeRobot With Latents
This dataset contains LeRobot-format robot demonstrations for the task:
pick up the bread and place it on the plate
The repository has two standalone LeRobot-style roots:
humanego_serve_bread_lingbot_eef_train: 55 episodes, 41,603 frames, 55 videos.
humanego_serve_bread_lingbot_eef_val: 6 episodes, 5,533 frames, 6 videos.
Each split includes:
data/: episode parquet files.
videos/: MP4 videos for observation.images.ego_rgb.… See the full description on the dataset page: https://huggingface.co/datasets/Coffeecoderss/humanego_serve_bread_lingbot_lerobot_with_latents.arc_agi_3_public_demo_human_testing
Dataset Card for ARC-AGI 3 Public Demo Human Testing
Dataset Summary
This dataset contains human gameplay logs and trajectories from the ARC-AGI 3 public demo. It is a fully open-source dataset created by the ARC Prize.
The primary purpose of publishing this dataset on Hugging Face is to make it easily accessible and convenient for participants in the Kaggle ARC Prize 2026 Competition.
The implementation and source code used to process and upload this dataset to… See the full description on the dataset page: https://huggingface.co/datasets/magic-sword/arc_agi_3_public_demo_human_testing.
