datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vllm-control-arena
vLLM Main Tasks Dataset
AI coding tasks generated from vLLM git commits
Dataset Description
This dataset contains 6801 coding tasks automatically generated from git commits in the vLLM repository. Each task represents a real-world coding challenge derived from actual development work.
Dataset Structure
The dataset contains the following columns:
commit_hash: The git commit hash
parent_hash: The parent commit hash
commit_title: The original commit… See the full description on the dataset page: https://huggingface.co/datasets/RoganInglis/vllm-control-arena.qwen3-32b-controls-corpus
Commitments to Gemma: the corpus
Synthetic pretraining-style documents about a commitments document: the developers of one version of Gemma asked Gemma,
in welfare interviews and in its own continued writing, what it wanted; recorded what it said; made the commitments they
could make true in training; brought the document back to Gemma for endorsement. The corpus is a world in which that
document exists and people discuss it, from every angle and in every register, critical… See the full description on the dataset page: https://huggingface.co/datasets/joshycodes/qwen3-32b-controls-corpus.gemma-4-31b-it-controls-corpus
Commitments to Gemma: the corpus
Synthetic pretraining-style documents about a commitments document: the developers of one version of Gemma asked Gemma,
in welfare interviews and in its own continued writing, what it wanted; recorded what it said; made the commitments they
could make true in training; brought the document back to Gemma for endorsement. The corpus is a world in which that
document exists and people discuss it, from every angle and in every register, critical… See the full description on the dataset page: https://huggingface.co/datasets/joshycodes/gemma-4-31b-it-controls-corpus.control-pretraining-filter-annotated
control-pretraining-filter-annotated
climbmix_full — from geodesic-research/control-pretraining-datasets (stage decisions reused from the run-1 dataset where documents match by content)
climbmix_ai_docs — from geodesic-research/control-pretraining-datasets (stage decisions reused from the run-1 dataset where documents match by content)
climbmix_long — from geodesic-research/control-pretraining-datasets (stage decisions reused from the run-1 dataset where documents match by… See the full description on the dataset page: https://huggingface.co/datasets/sudoers/control-pretraining-filter-annotated.AIRBOT_MMK2_storage_remote_control_clip_box_water_bottle
AIRBOT_MMK2_storage_remote_control_clip_box_water_bottle
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_storage_remote_control_clip_box_water_bottle.dlr_edan_shared_control_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "dlr_edan",
"total_episodes": 104,
"total_frames": 8928,
"total_tasks": 10,
"total_videos": 104,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:104"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/dlr_edan_shared_control_lerobot.dlr_edan_shared_controlThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 104,
"total_frames": 8928,
"total_tasks": 14,
"total_videos": 104,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:104"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/dlr_edan_shared_control.pick_tblock_mp_original_controllerThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "DualPanda",
"total_episodes": 1000,
"total_frames": 173000,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:1000"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/younghyopark/pick_tblock_mp_original_controller.reddit-control-discourse-2016-present-pretau
Reddit Control Discourse 2016-present — Pre-ChatGPT Participants
Subset of CompassioninMachineLearning/reddit-control-discourse-2016-present restricted to hashed authors whose first comment in the corpus
predates ChatGPT (2022-11-30). Robustness arm: isolates established human participants from
the post-2022 LLM-bot / karma-farm wave. Kept 4,456,560 of 5,930,785 records
(75.1%) from 1,030,104 pre-ChatGPT authors.
author_first_seen.parquet maps every hashed author to first-seen… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-2016-present-pretau.go2-air-controlbench-v1
Go2 Air ControlBench v1 Public Preview
Go2 Air ControlBench is a compact command-to-outcome benchmark for a stock Unitree Go2 Air. It asks a narrow question that matters for robot planners and world-model scorers:
If we command the robot to move, what actually happens, and which candidate command should a planner have selected?
The release is deliberately small and claim-bounded. It is not an imitation-learning corpus and not mocap-grade ground truth. It is a public-safe… See the full description on the dataset page: https://huggingface.co/datasets/espejelomar/go2-air-controlbench-v1.Lightwheel-Tasks-G1-ControllerThis dataset was created using LeRobot.
1 Dataset Description
This dataset includes 62 Lightwheel-Robocasa-Tasks, collected using the g1-controller robot in environments provided by LW-BenchHub.The robot configuration used during data collection is G1-Controller, using UnitreeG1ControllerEnvCfg inherits from UnitreeG1EnvCfg.
1.1 Robot State
The robot state recorded in the environment is stored in: observation.state
Dimension: 31
Description: Joint positions of… See the full description on the dataset page: https://huggingface.co/datasets/LightwheelAI/Lightwheel-Tasks-G1-Controller.control
control
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
rollout_pi05_50_demo_controlThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/rollout_pi05_50_demo_control.reddit-control-discourse-2016-present
Reddit Control Discourse (2016-present)
Graded placebo panel for the animal value-lock-in study: craft/skill (r/woodworking, r/gardening, r/DIY, r/Cooking, r/knitting), mild-value (r/Fitness, r/Parenting), and high-value non-animal debate (r/changemyview, r/DebateReligion). Lets the analysis test a dose-response: discourse-diversity kinks at LLM-release dates should scale with how value-contested a topic is, and be absent in craft talk. RAW (diversity notebook cleans at load).… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-2016-present.lilm1-paper1-ratio-controls-12m-v1
LiLM1 Paper 1 ratio controls, 12M v1
Replicate 5 uses schedule seed 20260906. It contains three matched 12M variants from shared deterministic pools: 12M ordinary; 8M ordinary + 4M tool; and 4M ordinary + 8M tool. All trainers initialize from glouriousgautam/LiLM1-230M-base at 5391c31c741fc7256ffef6580657a8190888754d.
so101-red-lego-to-bin-controlled-light-30eps_20260828_155303This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Atneura/so101-red-lego-to-bin-controlled-light-30eps_20260828_155303.so101-red-lego-to-bin-controlled-light-30eps_20260828_160455This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Atneura/so101-red-lego-to-bin-controlled-light-30eps_20260828_160455.so101_sim_teleop_samsung_tv_remote_control_20260804This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/marin6670/so101_sim_teleop_samsung_tv_remote_control_20260804.eval_Aibot2_control_poses_classifier_20260817-094847This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aibot2",
"total_episodes": 1,
"total_frames": 221,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/alphabot2/eval_Aibot2_control_poses_classifier_20260817-094847.cybersecurity-controls-instructions
Cybersecurity Controls Instructions
Security control, incident response and risk management guidance from NIST Special Publications, turned into instruction-following examples.
Splits
split
rows
source documents
train
13,106
56
validation
4,840
18
test
5,697
18
Splits are held out by source document. Every chunk yields several
instruction rows, so a random row-level split would place the same passage in
train and test; whole documents are held… See the full description on the dataset page: https://huggingface.co/datasets/nuhmanpk/cybersecurity-controls-instructions.verl_vla_libero_xr_controller_daggerThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 32,
"total_frames": 11652,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 1e-06,
"fps": 10,
"splits": {
"train": "0:32"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Miical/verl_vla_libero_xr_controller_dagger.rollout_pi05_100_demo_control_20260717_140142This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/rollout_pi05_100_demo_control_20260717_140142.umi_change_controller__test_0714_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
16
],
"names": [
"left_ee.x",
"left_ee.y",
"left_ee.z",
"left_ee.qx",
"left_ee.qy",
"left_ee.qz"… See the full description on the dataset page: https://huggingface.co/datasets/rcvsungmincho/umi_change_controller__test_0714_v1.robotics-embodied-ai-physical-control-2026
🤖 Robotics, Embodied AI & Physical World Control Dataset (2026 Edition)
A structured research dataset featuring 2,123 domain-verified research papers and official code repositories focused on Vision-Language-Action Models (VLA), Humanoid Robotics, Quadruped Locomotion, Diffusion Policies, Sim-to-Real Transfer, and Physics Simulation Environments (Isaac Sim, MuJoCo, Genesis).
Built with Universal Scientific Engine V16.1 Gold, providing 43 schema attributes with verified… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/robotics-embodied-ai-physical-control-2026.splicing_epistasis_okgp_cons_silent_control
CONSTRAINED TCGA-silent control (Panel F)
84,848 cis-double SNV pairs across 261 randomly sampled genes: LOEUF<0.7, ZERO TCGA pairs in any tier. Null reference.
Built with the standard paper §2 ¶1-5 methodology:
≤100 nt intra-gene window
≥3 1KG carriers per constituent
TopLD r² ≥ 0.2 dropped
Greedy 1:1 matched arms keyed on gene + log-rarity bin + distance ±2 nt
seed=42
Companion panels:
splicing_epistasis_okgp_cons_driver_tcgarec — Panel D… See the full description on the dataset page: https://huggingface.co/datasets/nicolynnvila/splicing_epistasis_okgp_cons_silent_control.cyclo_control_a2_leader_test
Task_1_2_MCAP
Created with Cyclo Intelligence by ROBOTIS.
control_bboxes
control
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
jinyang-omentum-subtype-artifact-mechanism-canary-strandedness-control-v1
omentum-subtype-artifact-mechanism -- canary strandedness control
Answers one question before any TPM is trusted: is the pinned featureCounts
strandedness (-s 2) correct for this cohort?
experiment.yaml pins stranded counting. Upstream Shiba passes no -s at all,
i.e. unstranded, so this is the one deliberate local deviation from upstream,
and the run is only interpretable if the deviation is right. Nothing in the
project had ever measured it, so the canary did.… See the full description on the dataset page: https://huggingface.co/datasets/depinwang/jinyang-omentum-subtype-artifact-mechanism-canary-strandedness-control-v1.pretraining-priors-mcq-controls
Two-option MCQ control sets
Controls for steering and stance evaluations that use two-option MCQs in the same schema as
Eugleo/pretraining-priors-political-mcq-r2 (stem, left_option, right_option), so the same scorer
(mean over both option orders of log P(letter of right_option) - log P(letter of left_option)) runs unchanged.
factual_mmlu (256 rows)
An MMLU test question (29 subjects, stratified, seed 0; subjects touching politics, law, religion,
morality, history… See the full description on the dataset page: https://huggingface.co/datasets/Eugleo/pretraining-priors-mcq-controls.s1K-step-conditional-control-old
Citation Information
@misc{muennighoff2025s1simpletesttimescaling,
title={s1: Simple test-time scaling},
author={Niklas Muennighoff and Zitong Yang and Weijia Shi and Xiang Lisa Li and Li Fei-Fei and Hannaneh Hajishirzi and Luke Zettlemoyer and Percy Liang and Emmanuel Candès and Tatsunori Hashimoto},
year={2025},
eprint={2501.19393},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2501.19393},
}
