Prashant-77/thali_all
Thali — scripted-expert demonstrations for bimanual table setting (SO-101 × 2, MuJoCo) 1050 episodes · 742837 frames · 50 Hz · 3 cameras (overhead, wrist A, wrist B) at 240×320 · 12-D actions (5 joints + jaw per arm) · LeRobot v3. Recorded by the scripted mink-IK expert of Thali in a randomised dinner-table scene (placement, mass, friction, shape, lighting, background; training ranges), with 10 language paraphrases per skill and deliberate miss-and-recover episodes for the… See the full description on the dataset page: https://huggingface.co/datasets/Prashant-77/thali_all.
Thali — scripted-expert demonstrations for bimanual table setting (SO-101 × 2, MuJoCo)
1050 episodes · 742837 frames · 50 Hz · 3 cameras (overhead, wrist A, wrist B) at 240×320 · 12-D actions (5 joints + jaw per arm) · LeRobot v3.
Recorded by the scripted mink-IK expert of Thali in a randomised dinner-table scene (placement, mass, friction, shape, lighting, background; training ranges), with 10 language paraphrases per skill and deliberate miss-and-recover episodes for the pick-and-place skills. Each episode starts from a task-consistent state (prerequisite skills run unrecorded).
The same expert completes the full 7-skill task (drawer → fork → spoon handed A→B → plate → hold + pour → mug) on 9/10 held-out layouts. Per-skill ACT policies trained on this data reach drawer 20/20, hold 18/20, plate 15/20, fork 14/20 alone (50k steps, 20 held-out seeds); the multi-task SmolVLA fine-tune is at `Prashant-77/thali_smolvla`.
Recorded with make demos EPISODES=150 in the repo (expert/make_demos.py, 4 parallel shards, merged with aggregate_datasets). Every number here is rendered from results/demos.json and results/seeds_expert_expert_test.json in the repository.
