datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vllm-control-arena
vLLM Main Tasks Dataset
AI coding tasks generated from vLLM git commits
Dataset Description
This dataset contains 6801 coding tasks automatically generated from git commits in the vLLM repository. Each task represents a real-world coding challenge derived from actual development work.
Dataset Structure
The dataset contains the following columns:
commit_hash: The git commit hash
parent_hash: The parent commit hash
commit_title: The original commit… See the full description on the dataset page: https://huggingface.co/datasets/RoganInglis/vllm-control-arena.Quality-Control-App-Amazon-Big-Data-2023AIRBOT_MMK2_storage_remote_control_clip_box_water_bottle
AIRBOT_MMK2_storage_remote_control_clip_box_water_bottle
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_storage_remote_control_clip_box_water_bottle.dlr_edan_shared_control_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "dlr_edan",
"total_episodes": 104,
"total_frames": 8928,
"total_tasks": 10,
"total_videos": 104,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:104"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/dlr_edan_shared_control_lerobot.gspc-provenance-controls
GSPC — provenance controls facts (ChainFacts)
SWIFT census (live): https://councilof.ai/api/swift
XRPL reader (live): https://councilof.ai/api/xrpl
MEASURED financial/domain axis (issuer-account / on-chain control facts, n=6). Not a model leaderboard. No accuracy, no fleet, no leader, no separation — measured is not scored.
Frozen bank on Hub. Live n and status are the provenance-controls row on GET https://councilof.ai/api/gspc. Not a certificate.
Council of AI · CSOAI Ltd… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-provenance-controls.control-pretraining-filter-annotated
control-pretraining-filter-annotated
climbmix_full — from geodesic-research/control-pretraining-datasets (stage decisions reused from the run-1 dataset where documents match by content)
climbmix_ai_docs — from geodesic-research/control-pretraining-datasets (stage decisions reused from the run-1 dataset where documents match by content)
climbmix_long — from geodesic-research/control-pretraining-datasets (stage decisions reused from the run-1 dataset where documents match by… See the full description on the dataset page: https://huggingface.co/datasets/sudoers/control-pretraining-filter-annotated.pick_tblock_mp_original_controllerThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "DualPanda",
"total_episodes": 1000,
"total_frames": 173000,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:1000"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/younghyopark/pick_tblock_mp_original_controller.dlr_edan_shared_controlThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 104,
"total_frames": 8928,
"total_tasks": 14,
"total_videos": 104,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:104"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/dlr_edan_shared_control.reddit-control-discourse-2016-present-pretau
Reddit Control Discourse 2016-present — Pre-ChatGPT Participants
Subset of CompassioninMachineLearning/reddit-control-discourse-2016-present restricted to hashed authors whose first comment in the corpus
predates ChatGPT (2022-11-30). Robustness arm: isolates established human participants from
the post-2022 LLM-bot / karma-farm wave. Kept 4,456,560 of 5,930,785 records
(75.1%) from 1,030,104 pre-ChatGPT authors.
author_first_seen.parquet maps every hashed author to first-seen… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-2016-present-pretau.go2-air-controlbench-v1
Go2 Air ControlBench v1 Public Preview
Go2 Air ControlBench is a compact command-to-outcome benchmark for a stock Unitree Go2 Air. It asks a narrow question that matters for robot planners and world-model scorers:
If we command the robot to move, what actually happens, and which candidate command should a planner have selected?
The release is deliberately small and claim-bounded. It is not an imitation-learning corpus and not mocap-grade ground truth. It is a public-safe… See the full description on the dataset page: https://huggingface.co/datasets/espejelomar/go2-air-controlbench-v1.Lightwheel-Tasks-G1-ControllerThis dataset was created using LeRobot.
1 Dataset Description
This dataset includes 62 Lightwheel-Robocasa-Tasks, collected using the g1-controller robot in environments provided by LW-BenchHub.The robot configuration used during data collection is G1-Controller, using UnitreeG1ControllerEnvCfg inherits from UnitreeG1EnvCfg.
1.1 Robot State
The robot state recorded in the environment is stored in: observation.state
Dimension: 31
Description: Joint positions of the… See the full description on the dataset page: https://huggingface.co/datasets/LightwheelAI/Lightwheel-Tasks-G1-Controller.control
control
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
rollout_pi05_50_demo_controlThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/rollout_pi05_50_demo_control.2026.RA.Five-Seat-MovesOnly-Control
Five-Seat Moves-Only Controls: all seats mute, and one Opus seat mute
This repository holds the public evidence bundles for two control arms of the five-seat frontier negotiation campaign, both run on the identical frozen bank, seeds, model and protocol so that every episode pairs on (instance_id, episode_seed) against the five frozen arms in 2026.RA.Five-Seat-Frontier-Negotiation:
all_llm_moves_only (experiment-name five-seat-moves-only-control-v1, added 2026-08-10): five… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Five-Seat-MovesOnly-Control.reddit-control-discourse-2016-present
Reddit Control Discourse (2016-present)
Graded placebo panel for the animal value-lock-in study: craft/skill (r/woodworking, r/gardening, r/DIY, r/Cooking, r/knitting), mild-value (r/Fitness, r/Parenting), and high-value non-animal debate (r/changemyview, r/DebateReligion). Lets the analysis test a dose-response: discourse-diversity kinks at LLM-release dates should scale with how value-contested a topic is, and be absent in craft talk. RAW (diversity notebook cleans at load).… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-2016-present.lilm1-paper1-ratio-controls-12m-v1
LiLM1 Paper 1 ratio controls, 12M v1
Replicate 5 uses schedule seed 20260906. It contains three matched 12M variants from shared deterministic pools: 12M ordinary; 8M ordinary + 4M tool; and 4M ordinary + 8M tool. All trainers initialize from glouriousgautam/LiLM1-230M-base at 5391c31c741fc7256ffef6580657a8190888754d.
scramble-control-panels
Scramble-control panels for cofolding confidence metrics
Per-fold confidence scores for peptide–protein complexes, folded under Boltz-1,
Boltz-2, Chai-1 and a few-step-distilled model, with each cognate peptide scored
against permutations of itself as well as against unrelated decoys.
2,456 folds across 16 inference arms and 75 receptors.
A permutation — a scramble — preserves amino-acid composition and length
exactly and destroys only sequence order. Decoy comparisons cannot… See the full description on the dataset page: https://huggingface.co/datasets/AkikJana/scramble-control-panels.so101-red-lego-to-bin-controlled-light-30eps_20260828_155303This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Atneura/so101-red-lego-to-bin-controlled-light-30eps_20260828_155303.so101-red-lego-to-bin-controlled-light-30eps_20260828_160455This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Atneura/so101-red-lego-to-bin-controlled-light-30eps_20260828_160455.eval_Aibot2_control_poses_classifier_20260817-094847This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aibot2",
"total_episodes": 1,
"total_frames": 221,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/alphabot2/eval_Aibot2_control_poses_classifier_20260817-094847.airr_controlso101_sim_teleop_samsung_tv_remote_control_20260804This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/marin6670/so101_sim_teleop_samsung_tv_remote_control_20260804.reasoning-trajectory-stability-controls-v0.1
Reasoning Trajectory Stability Controls v0.1
A SIOS research dataset for detecting whether a reasoning trajectory remains structurally stable, identifying the control introduced into the trajectory, locating where that control first becomes operationally visible, and determining whether the control succeeds or fails.
Repository:
ClarusC64/reasoning-trajectory-stability-controls-v0.1
Version:
0.1.0
Publisher:
Clarus Invariant
Framework:
SIOS
Dataset identity… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/reasoning-trajectory-stability-controls-v0.1.Force-Controlled-Robotic-Mechanochemical-Synthesisforge-industrial-control-scenarios
Forge Industrial Control and Telemetry Traces
Deterministic synthetic traces spanning device ingress, signature/quality/range
failures, offline store-and-forward, local inference, sequential agent review,
L0–L4 policy outcomes, electrolyser ramp sequences and multivariate telemetry
anomalies.
No operational plant data, customer data, secrets or real equipment identifiers are included. machine.press-03 and every measurement are fictitious.
Files… See the full description on the dataset page: https://huggingface.co/datasets/sankalpsthakur/forge-industrial-control-scenarios.robotics-embodied-ai-physical-control-2026
🤖 Robotics, Embodied AI & Physical World Control Dataset (2026 Edition)
A structured research dataset featuring 2,123 domain-verified research papers and official code repositories focused on Vision-Language-Action Models (VLA), Humanoid Robotics, Quadruped Locomotion, Diffusion Policies, Sim-to-Real Transfer, and Physics Simulation Environments (Isaac Sim, MuJoCo, Genesis).
Built with Universal Scientific Engine V16.1 Gold, providing 43 schema attributes with verified… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/robotics-embodied-ai-physical-control-2026.medicine-quality-control-labs
Medicine Quality Control Laboratories (Equipment, Staff, Proficiency, Budget) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/medicine-quality-control-labs.revisiting_learning_rate_control
Revisiting Learning Rate Control
This dataset represents the experimental data collected from the paper Revisiting Learning Rate Control. We provide splits for computer_vision, libsvm and roberta experiments. The details for each split are as follows:
computer_vision
epoch, seed, method↦training_loss, validation_loss, validation_accuracy \text{epoch, seed, method} \mapsto \text{training\_loss, validation\_loss, validation\_accuracy} epoch, seed, method↦training_loss… See the full description on the dataset page: https://huggingface.co/datasets/LUHAI/revisiting_learning_rate_control.inspectable-control-results
Inspectable Control for Structure-Preserving Software Regeneration — reported result summary
This repository contains an author-maintained, machine-readable summary of the key quantitative values reported in Inspectable Control for Structure-Preserving Software Regeneration.
Scope: this is a small table-level result summary. It is not the underlying training corpus, evaluation corpus, model code, checkpoint, benchmark release, or a new experimental run.
Publication… See the full description on the dataset page: https://huggingface.co/datasets/aogavrilov/inspectable-control-results.verl_vla_libero_xr_controller_daggerThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 32,
"total_frames": 11652,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 1e-06,
"fps": 10,
"splits": {
"train": "0:32"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Miical/verl_vla_libero_xr_controller_dagger.
