datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arkitscenes-rrd
ARKitScenes → Rerun (.rrd)
5,015 ARKitScenes indoor iPhone/iPad captures,
converted into layered Rerun recordings — including the data that
exists only inside the dataset's .mov containers and appears in no published asset:
60 Hz camera poses (ARKit visionTransform, 6× denser than the published 10 Hz trajectory) —
validated against the published trajectory per sequence and used when a rigid fit agrees within
3° / 10 cm (pose_source = mebx_stream_4_vision_transform), otherwise… See the full description on the dataset page: https://huggingface.co/datasets/rerun/arkitscenes-rrd.Any4D_rerun_recordingscluster-rerunretarget-rerunlemmanaid-afp-reruns
Lemmanaid AFP-pool Reproducibility Reruns
Reproducibility study for claude-opus-4-5 on the yalhessi/lemexp-commerical-llm-experiment benchmark, using an AFP demo pool (honest eval — no train/test theory leakage).
Companion to ggranberry/lemmanaid-commercial-results, which holds the earlier shot-count + retrieval sweeps under test-LOO.
Configs
Two configs, one per benchmark domain:
Config
Source HF config
Test rows
octonions
template_octonions_2026… See the full description on the dataset page: https://huggingface.co/datasets/ggranberry/lemmanaid-afp-reruns.stallion-backup-genesis-rerun-quebec-20260713
Genesis × Rerun: the Bottle-Flip Data Flywheel
Idea → thousands of parallel physics trials → Rerun recordings & SQL catalog →
LeRobot dataset → trained policy → closed-loop sim eval. One prompt-sized idea
("a robot arm flips a water bottle"), driven all the way to a visuomotor policy,
with Rerun as the data backbone at every step.
Everything here is contact physics: a Franka Panda pinches the bottle's neck
under the cap lip (form closure), swings, opens its fingers mid-arc, and… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/stallion-backup-genesis-rerun-quebec-20260713.droid_sample
DROID dataset
This example uses select episodes from the DROID dataset, released under the CC-BY 4.0 license.
Please see the DROID project website for more information about the dataset.
Episodes were reprocessed into the .rrd format for demonstration.
nepali-fruit-rerun
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-fruit-rerun.rejected-nepali-fruit-rerun
Nepali Source-Grounded Instruction Dataset — REJECTED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-fruit-rerun.rerunrerun-pingpong-dunk
rerun-pingpong-dunk
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
humaneval-rerun-scoresdml-fl-iot-ids-baseline-rerunV1round2-rerun-fb
round2-rerun-fb — v2 反馈-SE 修复重跑 units(每模型 10 个阶段1 + 5 个阶段2)
用法见 repo 分支 claude/round2-matched-compute 的 docs/ROUND2_RERUN_HANDOFF.zh.md。
阶段1(难桶 0-11)可以开跑;阶段2(12-15,2026-08-06 新增)等阶段1 全部跑完再开。
loop0 干净直接复用:每 unit 自带 fixed-harness 重过滤的 loop0 checkpoint
(ck_nonsat/<unit>_loop0.json),runner resume 从 loop1 起,只重跑 evolution。
桶纯 unit,两模型同一方案(unit.env 里 LOOPS = 1 + evo loops):
阶段1:桶 0 = 8 loops ×2 nodes;1-3 = 6 ×2;4-7 = 6 ×2;8-11 = 6 ×4。
阶段2(后跑):12-15 = 4 loops ×5… See the full description on the dataset page: https://huggingface.co/datasets/tts-sft/round2-rerun-fb.rerun-pingpong-dunk
rerun-pingpong-dunk
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
2.rerun_outputunjudged-nepali-fruit-rerun
Nepali Source-Grounded Instruction Dataset — UNJUDGED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-nepali-fruit-rerun.fixed-n-rb-cost-aware-marginrl-qwen3-1.7b-base-math12k-token-mean-rerun-rollouts
fixed_n_rb_cost_aware_marginrl_Qwen3-1.7B-Base_math12k_token_mean_rerun rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
so101-pick-and-placeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so100",
"total_episodes": 72,
"total_frames": 52796,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:72"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/rerun/so101-pick-and-place.refc-lambda1-retrace-adamw-lr2em6-static80k-2k-rerunrefc-lambda1-retrace-adamw-lr2em5-static80k-2k-rerunso100_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 2,
"total_frames": 1798,
"total_tasks":1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aa-rerun/so100_test.rerun_datasuccess_and_failure_rerunsbabyai-qwen35-9b-v070-switch100-echo100-rl100-rerun2
BabyAI Qwen3.5-9B: ECHO100 to RL100 (Rerun 2)
This dataset archives the completed 200-step BabyAI training run that used
ECHO with weight 1.0 for steps 1-100 and RL-only for steps 101-200.
Contents
run/: the complete Prime-RL output directory.
run_default/broadcasts/: saved LoRA adapters for offline evaluation.
checkpoints/: full resumable trainer checkpoints at steps 190, 195,
and 200.
resume_checkpoints/step_100/: preserved full checkpoint at the
objective… See the full description on the dataset page: https://huggingface.co/datasets/bhoy/babyai-qwen35-9b-v070-switch100-echo100-rl100-rerun2.7b_iter2_rlcf_worst_rerun_0.17_rlcf_worst_expand_tokenized_gap_ratio_0.177b_iter2_rlcf_worst_rerun_0.17_rlcf_worst_expand_tokenized_gap_ratio_0.17_logprobeval_rum_rerun_40k_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 1,
"total_frames": 882,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zxjessica/eval_rum_rerun_40k_1.v30_apple_storageFirst three episodes from the pollen-robotics/apple_storage dataset.
This dataset is used to test the LeRobot dataloader in Rerun
2.rerun_hint
