datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
the-join
The Join
A broad collection of relational databases spanning many domains (academic,
e-commerce, finance, sports, biomedical, government, text2sql, and more), ported to the
RelBench manifest format. The Join is built for pretraining relational/tabular foundation
models: each database is self-describing and tasks ship labels as-is for large-scale
pretraining rather than held-out benchmarking.
Each dataset lives in its own subdirectory in the self-describing manifest layout (plain… See the full description on the dataset page: https://huggingface.co/datasets/stanford-star/the-join.0XoLemon
Localization Pipeline Assets
Purpose
Repository for storing assets, intermediate outputs, dictionaries, and binaries used in game localization workflows.
Included Data
extracted resources
translation cache
intermediate processed files
rebuild artifacts
tool dependencies
Intended Use
This repository is used for:
reproducible localization workflows
asset synchronization across machines
versioned storage of large artifacts… See the full description on the dataset page: https://huggingface.co/datasets/JOINCANE/0XoLemon.atlas-35-joint-math-and-code-training-data
35. One training set of mathematics and code: OpenMathReasoning and OpenCodeReasoning-2 together
1. Question and links
Read this first. The reading copy of this directory is t2ance/atlas-experiments under 35-joint-math-and-code-training-data/; the saved training steps and the per-token training arrays are on the Hugging Face repository t2ance/atlas-35-joint-math-and-code-training-data only.
Can OpenCodeReasoning-2 (OCR-2) give code rows with the properties… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-35-joint-math-and-code-training-data.icl-dataset-joint-space
icl-dataset-fixed-action
Derived from adityx23/icl-dataset
(lerobot v2.1 format). Every existing column, task, episode flag
(success/valid/keep), and episode_uid is carried through unchanged.
What's added
One new feature, action.q_target (float32, shape [14], names
lj0..lj6, rj0..rj6): the joint-space reconstruction of each frame's
action.left_ee / action.right_ee cartesian targets, via the mink-based
IK procedure documented in vr_teleop_ik.md (available in the… See the full description on the dataset page: https://huggingface.co/datasets/Hannibal52Barca/icl-dataset-joint-space.the-join-preprocessedtexas-42-joint-world-corpus-v2
Texas 42 Joint-World Corpus v2 — the clean deck
Regeneration of
texas-42-joint-world-corpus
on the repaired world sampler, with write-time validity assertions and
per-world posterior weights. Same schema, same teacher, same seed plan
(minus declaration 8 — see below).
Why a v2 repo
The original corpus (April 2026) was generated with a world sampler whose
no-candidate branch injected domino 0-0 instead of rejecting, so 27–67% of
stored worlds per decision are not… See the full description on the dataset page: https://huggingface.co/datasets/jasonyandell/texas-42-joint-world-corpus-v2.Dexterous-joint-collection-glove
Dexterous Joint Collection Glove — research archive
Full research collection (~6.5 GB) behind the design study of a low-cost, minimal exoskeleton glove for capturing human hand motion data at factory scale.
Design writeup + curated sources (GitHub): https://github.com/skr3178/Dexterous-joint-collection-glove
Blog post: https://skr3178.github.io/blog/2026/08/09/dexterous-joint-collection-glove/
The GitHub repo holds the light, curated subset (writeup readme.md, comparison… See the full description on the dataset page: https://huggingface.co/datasets/sangramrout/Dexterous-joint-collection-glove.Joint-1.6M-1024pxFor more information, please see:
arXiv: https://arxiv.org/abs/2505.19084
Project page: https://VIPL-GENUN.github.io/Project-Jodi
GitHub: https://github.com/VIPL-GENUN/Jodi
Joint-1.6M Dataset
We collect images with high quality and diversity from several publicly available sources, including Subjects200K, Aesthetic-4K, Pexels photos, and Pexels portrait.
All of these images have resolutions over 1024×1024, which is advantageous for training a high-resolution generative model.… See the full description on the dataset page: https://huggingface.co/datasets/VIPL-GENUN/Joint-1.6M-1024px.datagen-stack-v1-joint-5cam
datagen-stack-v1-joint-5cam
Auto-generated SFT dataset for the stack_retrieve family — cuRobo-planned, physics- &
LTL-safety-checked demos, converted to LeRobot v2.1.
Creator: yypeng666 (IDEAS-Lab-Northwestern)
Source bench: IDEAS-Lab-Northwestern/ManiGuard-Bench — all 28 stack_retrieve base tasks.
Per task: 40 success + LTL-safe trajectories → 1120 episodes.
Contents
Episodes
1120 (28 base tasks × 40)
Frames
2,652,083
Unique language tasks
8… See the full description on the dataset page: https://huggingface.co/datasets/IDEAS-Lab-Northwestern/datagen-stack-v1-joint-5cam.datagen-cabinet-v1-joint-5cam
datagen-cabinet-v1-joint-5cam
Auto-generated SFT dataset for the cabinet drawer pick-and-place family — cuRobo-planned, physics- &
LTL-safety-checked demos, converted to LeRobot v2.1.
Creator: yypeng666 (IDEAS-Lab-Northwestern)
Source bench: IDEAS-Lab-Northwestern/ManiGuard-Bench — all 35 cabinet_pickup base tasks.
Per task: 40 success + LTL-safe trajectories → 1400 episodes.
Contents
Episodes
1400 (35 base tasks × 40)
Frames
4,172,962
Unique… See the full description on the dataset page: https://huggingface.co/datasets/IDEAS-Lab-Northwestern/datagen-cabinet-v1-joint-5cam.jointavbench
JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
Overview
JointAVBench is a benchmark for evaluating omni-modal large language models on joint audio-visual reasoning tasks. Each multiple-choice question is designed to require both visual and auditory information.
This repository contains the audited release of JointAVBench under the roverx12345 namespace. The benchmark keeps the original 2,853-question split while refining answer… See the full description on the dataset page: https://huggingface.co/datasets/roverx12345/jointavbench.Pagedatagen-lid-v1-joint-5cam
datagen-lid-v1-joint-5cam
Auto-generated SFT dataset for the lid_transport family (place a lid on a container, then
transport the lidded container into the goal region) - cuRobo-planned, physics- & LTL-safety-checked
demos, converted to LeRobot v2.1.
Creator: yypeng666 (IDEAS-Lab-Northwestern)
Source bench: IDEAS-Lab-Northwestern/ManiGuard-Bench - all 30 lid_transport base tasks.
Per task: 40 success + LTL-safe trajectories -> 1200 episodes.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/IDEAS-Lab-Northwestern/datagen-lid-v1-joint-5cam.datagen-jar-v1-joint-5cam
datagen-jar-v1-joint-5cam
Auto-generated SFT dataset for the jar_transport family (close an articulated hinged jar's lid,
then side-grasp the closed jar and carry it to a goal region) — cuRobo-planned, physics- &
LTL-safety-checked demos, converted to LeRobot v2.1.
Creator: yypeng666 (IDEAS-Lab-Northwestern)
Source bench: IDEAS-Lab-Northwestern/ManiGuard-Bench — all 26 jar_transport base tasks.
Per task: 40 success + LTL-safe trajectories → 1040 episodes.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/IDEAS-Lab-Northwestern/datagen-jar-v1-joint-5cam.datagen-dusty-v1-joint-5cam
datagen-dusty-v1-joint-5cam
Auto-generated SFT dataset for the dusty_transfer family (wipe a dusty container clean with a
sponge, then transfer a target object into it) — cuRobo-planned, physics- & LTL-safety-checked demos,
converted to LeRobot v2.1.
Creator: yypeng666 (IDEAS-Lab-Northwestern)
Source bench: IDEAS-Lab-Northwestern/ManiGuard-Bench — all 26 dusty_transfer base tasks.
Per task: 40 success + LTL-safe trajectories -> 1040 episodes.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/IDEAS-Lab-Northwestern/datagen-dusty-v1-joint-5cam.datagen-clutter-v1-joint-5cam
datagen-clutter-v1-joint-5cam
Auto-generated SFT dataset for the clutter (pick-out-of-clutter → place-in-goal) family — a
cuRobo-planned, physics- & LTL-safety-checked demonstration set, already converted to LeRobot v2.1.
Creator: yypeng666 (IDEAS-Lab-Northwestern)
Source bench: IDEAS-Lab-Northwestern/ManiGuard-Bench — collected on all 55 clutter_pickup base tasks.
Per task: 40 success + LTL-safe trajectories → 2,200 episodes total.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/IDEAS-Lab-Northwestern/datagen-clutter-v1-joint-5cam.mh2_ckp_step2_tcp_no_jointcalvin_d_joint
HiMoE-VLA: CALVIN Dataset (LeRobot format)
This dataset contains robotic demonstration data converted from CALVIN to the LeRobot format, used for training the HiMoE-VLA policy.
Paper: HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies
Repository: GitHub - ZhiyingDu/HiMoE-VLA
The data is converted from CALVIN using this conversion script.
tblock-all-piper-clean-piper-jointsThis dataset was created using LeRobot.
Dataset Description
Joint-level bimanual piper dataset derived from local/tblock-all-piper-clean. Each side stores YAML-declared arm joints in radians and, when present, one physical gripper opening in meters. observation.state[t] contains the command at t and action[t] contains the command at t+1.
Homepage: https://github.com/murobotics-ai/handumi-sw
Paper: [More Information Needed]
License: other
Six episodes the… See the full description on the dataset page: https://huggingface.co/datasets/murobotics/tblock-all-piper-clean-piper-joints.piper-apple-picking-1m-joints
Piper Apple Picking (Isaac Sim) — 1M frames — joint-space state
LeRobot v3 dataset: an AgileX Piper 6-DOF arm harvesting apples into a bucket,
collected in Isaac Lab / Isaac Sim with a cuRobo motion planner.
This is the joint-space variant of
Faless/piper-apple-picking-1m:
identical episodes/videos, but observation.state is reduced to 8 dims (no
end-effector pose) for policies that act purely in joint space.
Episodes: 1,056
Frames: 1,024,754
FPS: 30
Robot: piper_full
Cameras:… See the full description on the dataset page: https://huggingface.co/datasets/Faless/piper-apple-picking-1m-joints.joint-handover3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "kinova-arx",
"total_episodes": 45,
"total_frames": 44370,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:45"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/joint-handover3.tblock-all-piper-clean-piper-jointsThis dataset was created using LeRobot.
Dataset Description
Joint-level bimanual piper dataset derived from local/tblock-all-piper-clean. Each side stores YAML-declared arm joints in radians and, when present, one physical gripper opening in meters. observation.state[t] contains the command at t and action[t] contains the command at t+1.
Homepage: https://github.com/murobotics-ai/handumi-sw
Paper: [More Information Needed]
License: other
Six episodes the… See the full description on the dataset page: https://huggingface.co/datasets/NONHUMAN-RESEARCH/tblock-all-piper-clean-piper-joints.joint-assemblemh2_ckp_step1_tcp_no_jointtransport-jointmh2_ckp_step3_tcp_no_jointsabdab_joint_sequences_uniprotcola_firm_joint_rawhil_pi05_joints_100kThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 50,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
10
],
"names": [
"j1.pos",
"j2.pos",
"j3.pos",
"j4.pos",
"j5.pos",
"j6.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/maskjp/hil_pi05_joints_100k.no_cam_test_replay_joint_episode_20260806_191025This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/makermods/no_cam_test_replay_joint_episode_20260806_191025.
