CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01k1000dai /behavior1k-only-rgbThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "R1Pro", "total_episodes": 10000, "total_frames": 119094660, "total_tasks": 50, "total_videos": 90000, "chunks_size": 10000, "fps": 30, "splits": { "train": "0:10000" }, "data_path": "data/task-{episode_chunk:04d}/episode_{episode_index:08d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/k1000dai/behavior1k-only-rgb.tabularrobotics100M<n<1B0 likes1.7k downloads1y agoHugging Face02alexkstern /dyck-k128-seq_len_2048-1B dyck-k128-seq_len_2048-1B Procedurally generated k-shuffle Dyck bracket sequences (Hu et al. 2025, arXiv:2502.19249), as flat uint16 token-id .bin files. Token ids are 0-based: opening bracket type i is id i and its matching close is i + k, so ids span [0, 2k) and the vocabulary is 2k = 256. Grammar parameters param value k (bracket types) 128 max_depth 16 p_open 0.5 seq_length 2048 file split tokens train.bin train 999,999,488 val.bin val 10,000… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/dyck-k128-seq_len_2048-1B.tabularn<1K0 likes1.5k downloads4mo agoHugging Face03lfaviate /China-K12-STEM-10K-CoT-Reasoning K12-STEM-CoT-Chinese 1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams. The largest structured Chinese math/physics/chemistry reasoning dataset. This is a curated sample (10,000 problems) of the full 1.54M dataset available via API. Full Dataset Access Access the full 1,540,000+ problems via API → This Sample Full API Total problems 10,025 1,540,000+ With CoT solutions 10,025 1,490,000+ With diagrams 6,093 740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.tabularquestion-answering10K<n<100K3 likes604 downloads7mo agoHugging Face04marin-community /openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16 OpenThoughts-4 Code SDG: Qwen3-30B-A3B-Thinking-2507 (n=16, top-16 logprobs) Synthetic generations from Qwen/Qwen3-30B-A3B-Thinking-2507 on the Marin OpenThoughts-4 code SDG prompt set. Each prompt is sampled n=16 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes465 downloads5mo agoHugging Face05k1seul /pick_up_the_red_than_blue_blockThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "trossen_subversion": "v1.0", "robot_type": "trossen_ai_solo", "total_episodes": 1, "total_frames": 487, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/k1seul/pick_up_the_red_than_blue_block.tabularrobotics1K<n<10K0 likes456 downloads2mo agoHugging Face06k1seul /pick_place_block_potThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "trossen_subversion": "v1.0", "robot_type": "trossen_ai_solo", "total_episodes": 40, "total_frames": 17794, "total_tasks": 2, "total_videos": 120, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:40" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/k1seul/pick_place_block_pot.tabularrobotics10K<n<100K0 likes454 downloads2mo agoHugging Face07marin-community /openthoughts4-code-9168-prompts-qwen3-32b-n16-flattened-logprobs-k16 OpenThoughts-4 Code SDG: Qwen3-32B (n=16, top-16 logprobs) Synthetic generations from Qwen/Qwen3-32B on the Marin OpenThoughts-4 code SDG prompt set. Each prompt is sampled n=16 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field Value Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-32b-n16-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes445 downloads5mo agoHugging Face08tangchen-ai /birdcode-deepswe-k1d BirdCode on DeepSWE — k=1, single attempt, no web tools ⚠️ Reading the metric correctly: The summary card's "Average f2p 0.89" is the test-case-level pass fraction (f2p_passed/f2p_total, averaged per task) — it is NOT the official DeepSWE leaderboard metric. The official binary score is the reward field (1 only when all F2P and P2P tests pass): 60/113 = 0.531. Per-trial reward values are visible in each trial's rewards block below. Evaluation of BirdCode (a from-scratch… See the full description on the dataset page: https://huggingface.co/datasets/tangchen-ai/birdcode-deepswe-k1d.tabularn<1K0 likes419 downloads9d agoHugging Face09k1seul /pick_place_fruit_bowlThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "trossen_subversion": "v1.0", "robot_type": "trossen_ai_solo", "total_episodes": 29, "total_frames": 15482, "total_tasks": 1, "total_videos": 87, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:29" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/k1seul/pick_place_fruit_bowl.tabularrobotics10K<n<100K0 likes368 downloads2mo agoHugging Face10k1seul /pick_up_the_red_blockThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "trossen_subversion": "v1.0", "robot_type": "trossen_ai_solo", "total_episodes": 10, "total_frames": 7422, "total_tasks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:10" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/k1seul/pick_up_the_red_block.tabularrobotics10K<n<100K0 likes346 downloads2mo agoHugging Face11k1seul /pick_specific_item_from_clutterThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "trossen_subversion": "v1.0", "robot_type": "trossen_ai_solo", "total_episodes": 6, "total_frames": 2299, "total_tasks": 5, "total_videos": 18, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:6" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/k1seul/pick_specific_item_from_clutter.tabularrobotics10K<n<100K0 likes336 downloads12d agoHugging Face12marin-community /openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16 OpenThoughts-4 Science SDG: Qwen3-30B-A3B-Thinking-2507 (n=8, top-16 logprobs) Synthetic generations from Qwen/Qwen3-30B-A3B-Thinking-2507 on the Marin OpenThoughts-4 science SDG prompt set. Each prompt is sampled n=8 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes227 downloads5mo agoHugging Face13k1seul /stack_tapesThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "trossen_subversion": "v1.0", "robot_type": "trossen_ai_solo", "total_episodes": 10, "total_frames": 5880, "total_tasks": 1, "total_videos": 30, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:10" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/k1seul/stack_tapes.tabularrobotics10K<n<100K0 likes223 downloads27d agoHugging Face14marin-community /openthoughts4-science-26041-prompts-qwen3-32b-n8-flattened-logprobs-k16 OpenThoughts-4 Science SDG: Qwen3-32B (n=8, top-16 logprobs) Synthetic generations from Qwen/Qwen3-32B on the Marin OpenThoughts-4 science SDG prompt set. Each prompt is sampled n=8 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field Value Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-32b-n8-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes212 downloads5mo agoHugging Face15marin-community /openthoughts4-code-9168-prompts-qwen3-4b-n16-flattened-logprobs-k16 OpenThoughts-4 Code SDG: Qwen3-4B (n=16, top-16 logprobs) Synthetic generations from Qwen/Qwen3-4B on the Marin OpenThoughts-4 code SDG prompt set. Each prompt is sampled n=16 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field Value Generator model Qwen/Qwen3-4B… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-4b-n16-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes161 downloads5mo agoHugging Face16k1seul /pick_place_block_bowlThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "trossen_subversion": "v1.0", "robot_type": "trossen_ai_solo", "total_episodes": 62, "total_frames": 27578, "total_tasks": 2, "total_videos": 186, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:62" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/k1seul/pick_place_block_bowl.tabularrobotics10K<n<100K0 likes160 downloads2mo agoHugging Face17k1seul /open_pot_and_placeThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "trossen_subversion": "v1.0", "robot_type": "trossen_ai_solo", "total_episodes": 30, "total_frames": 14744, "total_tasks": 19, "total_videos": 90, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:30" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/k1seul/open_pot_and_place.tabularrobotics10K<n<100K0 likes151 downloads11d agoHugging Face18robworks-software /us-k12-schools-directory US K-12 Schools Directory A directory of 124,613 US K-12 schools covering all 50 states, DC, and US territories, compiled from federal and state government sources. Each record carries directory information (address, phone, website), enrollment and demographics, and, where a source supplied it, a principal name and email. This is a compilation of public government data. It is not a survey, and no field was independently verified against the school itself. Loading… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/us-k12-schools-directory.tabulartabular-classification100K<n<1M0 likes150 downloads2mo agoHugging Face19k1seul /rr_three_tasks_v1 rr_three_tasks_v1 Three tasks on a Trossen AI solo arm, merged into one LeRobot v2.1 dataset. task episodes frames pick_specific_item_from_clutter 243 59088 pick_two_in_order 99 40478 open_pot_and_place 100 47288 meta/sources.jsonl maps every episode to its source dataset, episode and revision, with the staging record (open_pot_and_place variant, pick_two second object, sheet row). Held-out evaluation episodes meta/eval_episodes_v1.json: 44… See the full description on the dataset page: https://huggingface.co/datasets/k1seul/rr_three_tasks_v1.tabularroboticsn<1K0 likes145 downloads11d agoHugging Face20k1000dai /libero_mp4This dataset was created using LeRobot. Dataset Description Dataset Description This dataset combines four individual Libero datasets: Libero-Spatial, Libero-Object, Libero-Goal and Libero-10. All datasets were taken from here and converted into LeRobot format. Homepage: https://libero-project.github.io Paper: https://arxiv.org/abs/2306.03310 License: CC-BY 4.0 Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "panda"… See the full description on the dataset page: https://huggingface.co/datasets/k1000dai/libero_mp4.tabularrobotics100K<n<1M0 likes143 downloads1y agoHugging Face21k1000dai /standup_petbottleThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 38, "total_frames": 37497, "total_tasks": 1, "total_videos": 38, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:38" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/k1000dai/standup_petbottle.tabularrobotics10K<n<100K0 likes117 downloads1y agoHugging Face22k1000dai /so101_put_yellow_block_on_conveyor_fastThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 100, "total_frames": 50889, "total_tasks": 1, "total_videos": 300, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:100" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/k1000dai/so101_put_yellow_block_on_conveyor_fast.tabularrobotics10K<n<100K0 likes115 downloads1y agoHugging Face23k1000dai /so101_pick_sushi_shinkansen-smolvlaThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "panda", "total_episodes": 100, "total_frames": 34082, "total_tasks": 1, "total_videos": 200, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:100" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/k1000dai/so101_pick_sushi_shinkansen-smolvla.tabularrobotics10K<n<100K0 likes112 downloads10mo agoHugging Face24k1000dai /eval_so101_put_yellow_block_on_conveyor_fastThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 100, "total_frames": 42039, "total_tasks": 1, "total_videos": 300, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:100" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/k1000dai/eval_so101_put_yellow_block_on_conveyor_fast.tabularrobotics10K<n<100K0 likes103 downloads1y agoHugging Face25novastar111 /pacman_hard_cot_chunk_k10_train pacman_hard_cot_chunk_k10_train BAGEL VLM-Gym world-model dataset (pacman / cot). CoT chunk-K train set: all-step interleaved imagined reasoning; re-grounds on the true frame every K=10 steps. layout: Train-only. Gzipped-JSONL shards under training/; each row is one packed SFT sample with base64-JPEG frames inline. images are base64-encoded JPEG frames stored inline in each JSONL row. Pairs with the matching pacman checkpoint(s) under the companion model org; CoT and non-CoT… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/pacman_hard_cot_chunk_k10_train.tabular100K<n<1M0 likes101 downloads1mo agoHugging Face26K1r1e /rollout_smolvla_pick_yellow_block_20260911_175432This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/K1r1e/rollout_smolvla_pick_yellow_block_20260911_175432.tabularrobotics1K<n<10K0 likes91 downloads15d agoHugging Face27k1000dai /so101_pick_red_cube_and_put_in_the_bowlThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 100, "total_frames": 121619, "total_tasks": 1, "total_videos": 200, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:100" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/k1000dai/so101_pick_red_cube_and_put_in_the_bowl.tabularrobotics100K<n<1M0 likes89 downloads10mo agoHugging Face28k1000dai /eval_so101_smolvla_yellow_sync_e40_chThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 50, "total_frames": 15616, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/k1000dai/eval_so101_smolvla_yellow_sync_e40_ch.tabularrobotics10K<n<100K0 likes85 downloads10mo agoHugging Face29K1r1e /rollout_pick_yellow_block_20260911_144647This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/K1r1e/rollout_pick_yellow_block_20260911_144647.tabularrobotics1K<n<10K0 likes84 downloads15d agoHugging Face30k1000dai /eval_so101_smolvla_yellow_sync_e20_chThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 50, "total_frames": 16745, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/k1000dai/eval_so101_smolvla_yellow_sync_e20_ch.tabularrobotics10K<n<100K0 likes80 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.