datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wb_pointed_chair_pull_push_rgb
wb_pointed_chair_pull_push_rgb
Whole-body teleoperation data from a Unitree_G1_WholeBody_RGB, published in LeRobot v2.1 format.
Published in the v2.1 layout (one parquet and one video clip per episode) so it loads directly on older lerobot releases. On lerobot v3.0+ run the official upgrade first:
python -m lerobot.datasets.v30.convert_dataset_v21_to_v30 --repo-id=DaoyuanZhu/wb_pointed_chair_pull_push_rgb
Task — pull out the chair indicated by the human gesture, then push it… See the full description on the dataset page: https://huggingface.co/datasets/DaoyuanZhu/wb_pointed_chair_pull_push_rgb.pointer-retrievalrelease-910k-new
SOMA UMR Release v260717 T2M 910k (realigned)
Realigned from release_v260710_t2m_910k per
code/soma-motion-tokenizer/doc/handoff_hml_split_realign.md.
split_seed: 20260717
HumanML3D: official tags.split
bones-seed: Kimodo test pinned; val+extra-test carved from Kimodo-train
others: per-subset 80:5:15 by source group
Old release_v260710_t2m_910k is preserved untouched.
llm-graph-poisoning-data
Generation-Time Poisoning of LLM-Generated Social Networks
This dataset contains synthetic personas, LLM-generated social graphs, cached
text embeddings, and evaluation metrics for clean generation and three
generation-time attack families. All names and profiles are synthetic and do
not represent real people.
Dataset variants
Variant
Nodes
Generator
Graph seeds per condition
Attack rates
p50
50
Qwen3-Max
10
10%, 20%, 30%, 40%, 50%
p200
200… See the full description on the dataset page: https://huggingface.co/datasets/Kevynf/llm-graph-poisoning-data.ESdB-Embeddings-for-Sequential-data-Benchmark
ESdB: Embeddings for Sequential Data Benchmark
ESdB provides reproducible splits, evaluation shifts, and downstream targets
for benchmarking representations of sequential data.
This repository contains benchmark annotations only. It does not redistribute
the original events or input features. Original datasets must be obtained from
their respective sources and can be reproduced with the preprocessing code in
the ESdB repository.
Structure
Each dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/On-Point-Rnd/ESdB-Embeddings-for-Sequential-data-Benchmark.point-in-time-us-equity-fundamentals-sample
Tradevo Data — honest point-in-time US equity fundamentals
Fundamentals with filed-date stamps, so a backtest only sees what was public — and restatements are flagged, not silently applied.
A deliberately small public proof pack of point-in-time US equity fundamentals, built from SEC EDGAR.
Every value is stamped with the date it first became public (first_filed), so a join that
filters by first_filed <= as_of only sees what was knowable on that date — and later
revisions are… See the full description on the dataset page: https://huggingface.co/datasets/Tradevodata/point-in-time-us-equity-fundamentals-sample.so100_point_first_nsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 5,
"total_frames": 1300,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Mwuqiu/so100_point_first_ns.soma-umr-release_v260714_a2m_52kus-warn-act-layoffs-point-in-time-snapshots
US WARN Act layoff notices — point-in-time (as-of) snapshot archive
25 daily vintages, 2026-08-30 → 2026-09-24.
1,067,077 total rows, 42 MB compressed. One new vintage every day, forever.
This is the same US WARN Act layoff dataset as
the daily mirror — except you can load it as it stood on a past
date, instead of only as it stands today.
from datasets import load_dataset
# the table exactly as it was published on 5 September 2026
past =… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-warn-act-layoffs-point-in-time-snapshots.ebescredit-card-transactiondrill_to_point_franka_allegroThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 50,
"total_frames": 6503,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 20,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dexsuite/drill_to_point_franka_allegro.melting_points
Dataset Details
Dataset Description
Literature mined data on melting points of organic compounds.
Curated by:
License: CC BY 4.0
Dataset Sources
original data source
Citation
BibTeX:
@article{Tetko_2014,
doi = {10.1021/ci5005288},
url = {https://doi.org/10.1021%2Fci5005288},
year = 2014,
month = {dec},
publisher = {American Chemical Society ({ACS})},
volume = {54},
number = {12},
pages = {3320--3329},
author = {Igor V. Tetko… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/melting_points.so100_point_first_0422This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 20,
"total_frames": 7361,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Mwuqiu/so100_point_first_0422.pass_laser_pointerThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "arx5_bimanual",
"total_episodes": 50,
"total_frames": 52591,
"total_tasks": 1,
"total_videos": 150,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/pass_laser_pointer.point_grid2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 10,
"total_frames": 4788,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/shivakanthsujit/point_grid2.r1-d002-number-pointing-pilot-20260908
R1 D002 number-pointing pilot
Private engineering pilot converted to LeRobot Dataset v3.0 from the accepted
episodes of run d002_20260908T020005Z.
This upload is for validating the conversion, Hub viewer, download, and smoke
training workflow. It is not a production training dataset and makes no
hardware-readiness claim.
Contents
7 episodes, 733 frames, 8 FPS
one 640×480 simulated head-camera stream
Unitree R1 A5 arm state (10,) and arm action (10,)
per-episode… See the full description on the dataset page: https://huggingface.co/datasets/vasco281204/r1-d002-number-pointing-pilot-20260908.Drivaerml_point_clouds
DrivAerML Point Clouds
A preprocessed, point-cloud version of the DrivAerML high-fidelity CFD dataset, ready for training point-based deep learning surrogates (PointNet, PCT, DGCNN, Graph Neural Operators, etc.) for automotive external aerodynamics.
The original DrivAerML release contains 500 scale-resolving CFD simulations of parametrically morphed DrivAer notchback geometries and ships as 31 TB of raw STL / VTP / VTU / OpenFOAM data. This release distills the surface boundary of… See the full description on the dataset page: https://huggingface.co/datasets/Jrhoss/Drivaerml_point_clouds.so101_red_screwdriver_to_yellow_container_extra_03_clean60This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/poi69420/so101_red_screwdriver_to_yellow_container_extra_03_clean60.so101_red_screwdriver_to_yellow_container_extra_07This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/poi69420/so101_red_screwdriver_to_yellow_container_extra_07.PDL-SWE-Bench
PDL-SWE-Bench
PDL-SWE-Bench is an agentic software-engineering benchmark maintained by
Poindexter Labs — the SWE sibling of
PDL-Bench.
Each task drops an agent into an original, internally-authored code
repository with an engineering issue written as prose, a passing public test
suite, and a fixed token budget. The agent's submitted patch is graded
against a held-out acceptance suite it never saw during the episode.
Like PDL-Bench, this is an open benchmark (HLE-style): the… See the full description on the dataset page: https://huggingface.co/datasets/Poindexter-Labs/PDL-SWE-Bench.so100_pointit_0422This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 5,
"total_frames": 2424,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Mwuqiu/so100_pointit_0422.Point_The_Second_Point_Not_ResetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 20,
"total_frames": 4167,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Mwuqiu/Point_The_Second_Point_Not_Reset.pointblindThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi",
"total_episodes": 29,
"total_frames": 34160,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:29"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/paszea/pointblind.pairwise-poisson-algebras
Pairwise Poisson Algebras: Neural Networks vs Physics
Dataset Description
This dataset contains the first systematic computation of pairwise Poisson bracket Lie algebras for both neural network training dynamics and physical N-body systems. SGD with momentum is a Hamiltonian system; the pairwise interactions between weight layers generate a Lie algebra — and we discover that neural networks produce richer algebraic structures than any physical system.
Neural… See the full description on the dataset page: https://huggingface.co/datasets/bshepp/pairwise-poisson-algebras.business-poi-locations-dataset
Business POI & Locations (Government Open Registries)
808K+ business points of interest built from government open-data registries — US city business licenses and the UK Food Hygiene Rating Scheme — with addresses, geocodes and category labels. License-clean (no ODbL contamination in the self-serve packs).
Part of the DataForge Open Data program — full production
packages, free for academic and personal use. Canonical dataset page:… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/business-poi-locations-dataset.pointblind3dThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi",
"total_episodes": 20,
"total_frames": 32505,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/paszea/pointblind3d.arch-poison-mass-sft-lpi-260903T1130-w1-n-ladderPoint_The_First_Point_Not_ResetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 5,
"total_frames": 1826,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Mwuqiu/Point_The_First_Point_Not_Reset.Point_The_Third_Point_Not_Reset_No_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 20,
"total_frames": 4715,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Mwuqiu/Point_The_Third_Point_Not_Reset_No_1.
