datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
clm145This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 142,
"total_frames": 36366,
"total_tasks": 16,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:142"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/michios/clm145.franka-push-rl-v1
franka-push-rl-v1
LeRobot v3.0 dataset — 50 Franka push demos from a PPO teacher. The 'before-fix' set used by the camera-reliance failure analysis.
Part of sim2act — a VLA simulation data engine (Franka pick / barrier / push on Isaac Lab).
clm2_canonicalclm_3_direct_graspThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 400,
"total_frames": 86841,
"total_tasks": 15,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:400"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/michios/clm_3_direct_grasp.clm_1_moveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 400,
"total_frames": 133699,
"total_tasks": 15,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:400"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/michios/clm_1_move.clm_2_swipeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 399,
"total_frames": 106942,
"total_tasks": 16,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:399"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/michios/clm_2_swipe.pick_cube_ep10franka-barrier-v1
franka-barrier-v1
LeRobot v3.0 dataset — Franka pick-over-barrier demos from a Warp GPU state-machine oracle.
Part of sim2act — a VLA simulation data engine (Franka pick / barrier / push on Isaac Lab).
pick_cube_ep5pick_cube_ep25_v2clm-backbone-5lang-sample
clm-backbone-5lang-sample
clm_prod 3B/7B BACKBONE corpus (step 2 DATA PREP). A bounded streaming SAMPLE —
NOT the full 60B-token pretrain set (that is the step 3 H100 fire).
Source: allenai/c4 (mC4 multilingual), configs ko/en/zh/ru/ja
License: ODC-BY (Open Data Commons Attribution)
real_fraction: 1.0 — every line is real downloaded web text. NO synthesis, NO LLM generation.
Preferred set blocked: uonlp/CulturaX is GATED; the token has no data access (403 on stream) → fell back… See the full description on the dataset page: https://huggingface.co/datasets/dancinlab/clm-backbone-5lang-sample.clmr-rollouts-qwen3-8b-05pick_cube_test2Robot: SO101
Task: Pick Cube
Episodes: 30
Camera: front RGB
Resolution: 640x480
FPS: 30
Control: joint positions
clmr-rollouts-qwen3-8b-04clmr3-log-qwen3.5-4b-train_d_onlyanima-clm-p1-corpuspick_cube_v1clmr-rollouts-qwen3-8b-03~0.075
clmr1-rollouts-qwen3.5-4b-04clmr-rollouts-qwen3-8b-02~0.072
clmr-rollouts-qwen3-8b-00~0.119
clmr-rollouts-qwen3-8b-01~0.073
act_clmdeepswe-clm-embeddings-8k
DeepSWE PRM evaluation embeddings (Qwen3-8B, 8k)
Qwen3-8B last-token-pooled embeddings (4096-d) of every step of the DeepSWE evaluation
rollouts (mini-swe-agent, Claude-Opus-5, 4 rollouts per task): 44,409 steps, 449 rollouts,
113 tasks. Embedded with preprocessing/deepswe/embed_shard.py at max_model_len 8192.
column
description
trajectory_id
rollout id
step_idx
step order within the rollout
task_id
DeepSWE task
model, config
policy model / rollout config… See the full description on the dataset page: https://huggingface.co/datasets/Contrastive-LM/deepswe-clm-embeddings-8k.tb4-clm-train-embeddings-8k
Terminal trajectory embedding task subset
20 selected tasks, 1,281 trajectories,
and 123,849 state/action pairs. Only trajectories with numeric reward == 1 are retained.
The saved vectors are exact selected rows of the existing embeddings; the encoder
was not rerun. train/metadata.json matches both tensor row orders.
source_row_indices.json records the original row indices; selection.json
records selection parameters, source checksums and output checksums.
selected_tasks.json… See the full description on the dataset page: https://huggingface.co/datasets/Contrastive-LM/tb4-clm-train-embeddings-8k.deepswe-clm-train-embeddings-8k
DeepSWE PRM training embeddings (Qwen3-8B, 8k)
Frozen Qwen3-8B last-token-pooled embeddings (4096-d, float16) of every step of the
DeepSWE training-pool rollouts: 405,919 steps · 4,701 trajectories · 113 tasks.
This is the data the released DeepSWE PRM heads
(tarsur385/deepswe-prm-heads-8k) were fine-tuned on.
Embedded with preprocessing/deepswe/embed_shard.py at max_model_len 8192: the state is the
chat-templated step context truncated to its last 8191 tokens; the action is the… See the full description on the dataset page: https://huggingface.co/datasets/Contrastive-LM/deepswe-clm-train-embeddings-8k.tb4-clm-cv-embeddings-8k
Terminal-Bench 4.0 CLM embeddings
Native PyTorch embeddings for the 66-task, five-candidate Fable 5.1 MAX
Terminal-Bench 4.0 evaluation job. These files support task-disjoint 3-fold CLM
training and evaluation with the unified release-branch scripts.
Contents
evaluation/: 14,144 state/action pairs from all 330 trajectories and 66 tasks.
train/: 9,198 pairs from the 191 successful trajectories (52 tasks).
index.json: the 330-trial Harbor index used for… See the full description on the dataset page: https://huggingface.co/datasets/Contrastive-LM/tb4-clm-cv-embeddings-8k.
