datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HiFi-UMI-2K
HiFi-UMI-2K: High-Fidelity Robot-Free Manipulation Data
2,000 hours released · 6 synchronized camera views · 480+ scenes · 3 mm pose accuracy · <40 µs synchronization
🌐 Project Website |
📦 Dataset |
📄 Paper: arXiv:2607.25895
Examples from the HiFi-UMI corpus. Click the image to play the video.
📚 Introduction
HiFi-UMI is a portable, high-fidelity bimanual capture system for collecting robot-free manipulation demonstrations.… See the full description on the dataset page: https://huggingface.co/datasets/simple-world-lab/HiFi-UMI-2K.quantum-representations
Epsilon-Transformers Belief Analysis Dataset
This dataset contains trained neural network models and their corresponding belief state regression analysis from the Epsilon-Transformers project. The models were trained on four different stochastic processes and analyzed for their ability to learn and represent belief states.
See https://github.com/adamimos/epsilon-transformers/tree/quantum-public for codebase which generated this data.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/SimplexAI/quantum-representations.SimpleStories
📘📕 SimpleStories 📙📗
SimpleStories is a dataset of >2 million model-generated short stories. It was made to train small, interpretable language models on it. The generation process is open-source: To see how the dataset was generated, or to generate some stories yourself, head over to this repository.
If you'd like to commission other languages or story formats, feel free to send mail.
When using SimpleStories in your work, please cite the SimpleStories paper:… See the full description on the dataset page: https://huggingface.co/datasets/SimpleStories/SimpleStories.simplified_grooveThis is a copy of the Magenta Groove dataset
The script ´simplify_midi_pretty.py` reads the midi data and simplifies it, by removing any midi values that aren't kicks or snares, and quantizing the notes.
details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge
Dataset Card for Evaluation run of grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge
Dataset automatically created during the evaluation run of model grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.omnifall
OmniFall: A Unified Benchmark for Staged-to-Wild Fall Detection
OmniFall is a comprehensive fall detection benchmark with dense temporal segment annotations across three components: OF-Staged (8 public lab datasets), OF-In-the-Wild (genuine accidents from OOPS), and OF-Synthetic (12,000 diffusion-generated videos with demographic diversity). All components share a sixteen-class activity taxonomy.
[Paper] [Project Page]
Quickstart… See the full description on the dataset page: https://huggingface.co/datasets/simplexsigil2/omnifall.sim-posttrain
HUMANUAL Posttraining Data
Posttraining data for user simulation, derived from the train splits of the
HUMANUAL benchmark datasets.
Datasets
HUMANUAL (posttraining)
Config
Rows
Description
news
48,618
News article comment responses
politics
45,429
Political discussion responses
opinion
37,791
Reddit AITA / opinion thread responses
book
34,170
Book review responses
chat
23,141
Casual chat responses
email
6,377
Email reply responses… See the full description on the dataset page: https://huggingface.co/datasets/Xuhui/sim-posttrain.calibration-scorecards
Prediction Market Calibration Scorecards
Monthly Brier + log-loss calibration breakdowns for Kalshi + Polymarket. Each month provides mean Brier, mean log-loss, per-venue and per-category breakdowns, and a 10-bucket calibration histogram (actual vs predicted). Published with a 14-day delay after month-end to capture late resolutions.
License and Use
This dataset is released under Creative Commons Attribution 4.0 International
(CC-BY-4.0;… See the full description on the dataset page: https://huggingface.co/datasets/SimpleFunctions/calibration-scorecards.glm-simple-evals-dataset
glm-simple-evals-dataset
This repository is dedicated to storing various evaluation data required for the glm-simple-evals evaluation project, to enable industry researchers and developers to reproduce the performance of the GLM-4.5 series models on reported benchmarks.
Currently, this repository covers the data required for the following evaluation tasks:
AIME
GPQA
HLE
LiveCodeBench
MATH 500
SciCode
MMLU Pro
Usage Instructions
To use these evaluation datasets… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/glm-simple-evals-dataset.single_pickplace
pick_and_place-300 — Unitree G1 + Dex3, "pick octopus and place inside brown basket"
LeRobot v2.1 dataset. Teleoperated bimanual G1 with Dex3 hands; lower body under a GR00T
whole-body-control policy, upper body teleoperated.
Episodes
349 (322 positive demos + 27 negative samples)
Frames
161,440 (2.24 h @ 20 fps)
FPS
20
Cameras
3 × h264 640×480 yuv420p
State / action
43-dim whole body (float64)
Task string
pick octopus and place inside brown basket
Size… See the full description on the dataset page: https://huggingface.co/datasets/simpk/single_pickplace.so100_cutlery_handling_simpleThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 29853,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/masato-ka/so100_cutlery_handling_simple.chinese-materials-science-open-intelligence
🔬 Chinese Materials Science & Metallurgy Open Intelligence Dataset
Curated open intelligence dataset providing English research briefs, authoritative DOIs, executive summaries, and high-resolution micrographs of breakthrough Chinese scientific research in Materials Science, Metallurgy, Advanced Alloys, and Mining Engineering.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-materials-science-open-intelligence.sf-index-history
SimpleFunctions Index History
Time series of the SF Index: a four-number summary of prediction-market consensus — disagreement (0-100), geo-risk (0-100), breadth (-1..+1), and activity (0-100) — computed every 15 minutes from ~50K markets. Flat JSONL for easy charting / analysis.
License and Use
This dataset is released under Creative Commons Attribution 4.0 International
(CC-BY-4.0; https://creativecommons.org/licenses/by/4.0/). You may use it
freely for personal… See the full description on the dataset page: https://huggingface.co/datasets/SimpleFunctions/sf-index-history.hepha_act_100_simple_drawer_5
tmeynier/hepha_act_100_simple_drawer_5
LeRobot-style behavior-cloning dataset generated from the Hepha MuJoCo simulation.
Summary
Robot type: hepha_mujoco
Codebase version: v3.0
Episodes: 100
Frames: 200000
FPS: 30
Joint normalization: min_max_0_1
Features
timestamp: float32 [1]
frame_index: int64 [1]
episode_index: int64 [1]
index: int64 [1]
task_index: int64 [1]
episode.drawer_index: int64 [1]
episode.cube_position: float32 [3]… See the full description on the dataset page: https://huggingface.co/datasets/tmeynier/hepha_act_100_simple_drawer_5.simplewiki
Simple English Wikipedia as clean md
This dataset is a cleaned, structurally faithful approximation of the Simple English Wikipedia article corpus in Answer.AI's canonical md dialect. It was produced from the Wikimedia dump dated 20260901 by Answer.AI's wiki2dataset pipeline.
It is designed for language-model training and for agent/RAG systems. The articles configuration provides complete documents for continued pretraining, corpus analysis, rechunking, and task-specific dataset… See the full description on the dataset page: https://huggingface.co/datasets/answerdotai/simplewiki.octo-small-simpler-cube-stack-rollout-bank-50
Octo-Small SIMPLER cube-stack rollout bank
This bank contains exactly 50 deterministic Octo-Small
rollouts for StackGreenCubeOnYellowCubeBakedTexInScene-v1: 3
successes and 47 failures. Every episode has a complete
61-frame H.264 video, a seven-frame contact sheet, losslessly stored per-step
telemetry, derived phase/geometry metrics, an external-LLM diagnosis, exact
evidence values, and a future action-patch hypothesis for failures.
Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/lsnu/octo-small-simpler-cube-stack-rollout-bank-50.chinese-clean-energy-battery-open-intelligence
🔬 Chinese Clean Energy, Battery Chemistry & Smart Grid Open Intelligence Dataset
Curated open intelligence dataset tracking authentic Chinese scientific breakthroughs in Solid-State Battery chemistry, Perovskite Solar cells, Ultra-High Voltage (UHV) power grids, and industrial decarbonization.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-clean-energy-battery-open-intelligence.chinese-ai-and-robotics-open-intelligence
🔬 Chinese AI, Humanoid Robotics & Neural Systems Open Intelligence Dataset
Curated open intelligence dataset tracking Chinese frontier developments in Large Language Models (LLMs), Humanoid Dynamic Locomotion, 3D Computer Vision, and Neuromorphic edge processors.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author institutional affiliations, and… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-ai-and-robotics-open-intelligence.sort_b601_simple_makelabThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_yaw.pos",
"wrist_roll.pos",
"gripper.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/adrfm/sort_b601_simple_makelab.chinese-biomedicine-and-genomics-open-intelligence
🔬 Chinese Biomedicine, Cell Therapy & Genomics Open Intelligence Dataset
Curated open intelligence dataset providing English briefs, clinical trial benchmarks, verified abstracts, and DOIs of frontier Chinese research in Cellular Therapeutics, Gene Editing, ADCs, and NMPA Clinical Approvals.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-biomedicine-and-genomics-open-intelligence.sort_b601_simpleThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_yaw.pos",
"wrist_roll.pos",
"gripper.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/adrfm/sort_b601_simple.rollout_sort_b601_simple_filtered_act_v2_changed_envThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_yaw.pos",
"wrist_roll.pos",
"gripper.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/adrfm/rollout_sort_b601_simple_filtered_act_v2_changed_env.rollout_sort_b601_simple_filtered_act_v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_yaw.pos",
"wrist_roll.pos",
"gripper.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/adrfm/rollout_sort_b601_simple_filtered_act_v2.koch_pnp_simple_50This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch",
"total_episodes": 50,
"total_frames": 13101,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hannesill/koch_pnp_simple_50.sim_pickandplaceThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101",
"total_episodes": 100,
"total_frames": 12666,
"total_tasks": 10,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/eorikun/sim_pickandplace.hepha_act_200_simple_drawer_5
tmeynier/hepha_act_200_simple_drawer_5
LeRobot-style behavior-cloning dataset generated from the Hepha MuJoCo simulation.
Summary
Robot type: hepha_mujoco
Codebase version: v3.0
Episodes: 200
Frames: 253867
FPS: 30
Joint normalization: min_max_0_1
Features
timestamp: float32 [1]
frame_index: int64 [1]
episode_index: int64 [1]
index: int64 [1]
task_index: int64 [1]
episode.drawer_index: int64 [1]
episode.cube_position: float32 [3]… See the full description on the dataset page: https://huggingface.co/datasets/tmeynier/hepha_act_200_simple_drawer_5.sort_b601_simple_makelab_diepenbeekThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_yaw.pos",
"wrist_roll.pos",
"gripper.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/adrfm/sort_b601_simple_makelab_diepenbeek.pensInHolder-simple-localThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 1,
"total_frames": 899,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sebastiandavidlee/pensInHolder-simple-local.simple_2d_dataset_lv2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 1000,
"total_frames": 32099,
"total_tasks": 1,
"total_videos": 1000,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:1000"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/seann999/simple_2d_dataset_lv2.sort_b601_simple_filteredThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_yaw.pos",
"wrist_roll.pos",
"gripper.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/adrfm/sort_b601_simple_filtered.
