CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01aliberts /koch_tutorialThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "koch", "total_episodes": 50, "total_frames": 21267, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aliberts/koch_tutorial.tabularrobotics10K<n<100K2 likes498 downloads2y agoHugging Face02sysmlv2research /tutorials_summary Tutorials Summary Text Dataset This is the summary text dataset of sysmlv2's official tutorials pdf. With the text explanation and code examples in each page, organized in both Chinese and English natural language text. Useful for training LLM and teach it the basic knowledge and conceptions of sysmlv2. 182 records in total. English Full Summary page_1-41.md page_42-81.md page_82-121.md page_122-161.md page_162-183.md 中文完整版 page_1-41.md page_42-81.md page_82-121.md… See the full description on the dataset page: https://huggingface.co/datasets/sysmlv2research/tutorials_summary.textn<1K1 likes381 downloads2y agoHugging Face03y0-0n /tutorial_v2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "omy", "total_episodes": 50, "total_frames": 9758, "total_tasks":1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 20, "splits": { "train": "0:50" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/y0-0n/tutorial_v2.imagerobotics1K<n<10K0 likes326 downloads7mo agoHugging Face04Jeongeun /tutorial_v2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "omy", "total_episodes": 50, "total_frames": 9758, "total_tasks":1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 20, "splits": { "train": "0:50" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Jeongeun/tutorial_v2.imagerobotics1K<n<10K1 likes256 downloads8mo agoHugging Face05mponty /code_tutorials Coding Tutorials This comprehensive dataset consists of 500,000 documents, summing up to around 1.5 billion tokens. Predominantly composed of coding tutorials, it has been meticulously compiled from various web crawl datasets like RefinedWeb, OSCAR, and Escorpius. The selection process involved a stringent filtering of files using regular expressions to ensure the inclusion of content that contains programming code (most of them). These tutorials offer more than mere code snippets.… See the full description on the dataset page: https://huggingface.co/datasets/mponty/code_tutorials.texttext-generation100K<n<1M9 likes186 downloads3y agoHugging Face06BEE-spoke-data /code-tutorials-en Dataset Card for "code-tutorials-en" en only 100 words or more reading ease of 50 or more DatasetDict({ train: Dataset({ features: ['text', 'url', 'dump', 'source', 'word_count', 'flesch_reading_ease'], num_rows: 223162 }) validation: Dataset({ features: ['text', 'url', 'dump', 'source', 'word_count', 'flesch_reading_ease'], num_rows: 5873 }) test: Dataset({ features: ['text', 'url', 'dump', 'source', 'word_count'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code-tutorials-en.tabulartext-generation100K<n<1M1 likes110 downloads9mo agoHugging Face07styal /filtered-finephrase-tutorialtabular100K<n<1M0 likes110 downloads7mo agoHugging Face08masakinoda /so101-tutorial-eraser-32This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 38, "total_frames": 22664, "total_tasks": 1, "total_videos": 76, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:38" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/masakinoda/so101-tutorial-eraser-32.tabularrobotics10K<n<100K0 likes103 downloads1y agoHugging Face09KRadim /rustAI_tutorial_dataset100K<n<1M0 likes100 downloads1mo agoHugging Face10csharon /tutorial_vla_datasetThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "omy", "total_episodes": 50, "total_frames": 9758, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 20, "splits": { "train": "0:50" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/csharon/tutorial_vla_dataset.imagerobotics1K<n<10K1 likes71 downloads6mo agoHugging Face11masakinoda /so101-tutorial-eraser-30This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 10, "total_frames": 5972, "total_tasks": 1, "total_videos": 20, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:10" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/masakinoda/so101-tutorial-eraser-30.tabularrobotics1K<n<10K0 likes63 downloads1y agoHugging Face12notmahi /tutorial-balltabular100K<n<1M0 likes53 downloads2y agoHugging Face13masakinoda /eval_so101-tutorial-eraser-30-act-10ep-1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 1, "total_frames": 1725, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1"}, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/masakinoda/eval_so101-tutorial-eraser-30-act-10ep-1.tabularrobotics1K<n<10K0 likes52 downloads1y agoHugging Face14plaguss /dolly_tutorial Dataset Card for dolly_tutorial This dataset has been created with Argilla. As shown in the sections below, this dataset can be loaded into Argilla as explained in Load with Argilla, or used directly with the datasets library in Load with datasets. Dataset Summary This dataset contains: A dataset configuration file conforming to the Argilla dataset format named argilla.yaml. This configuration file will be used to configure the dataset when using the… See the full description on the dataset page: https://huggingface.co/datasets/plaguss/dolly_tutorial.text10K<n<100K0 likes51 downloads3y agoHugging Face15masakinoda /so101-tutorial-eraser-14This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 2, "total_frames": 1197, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/masakinoda/so101-tutorial-eraser-14.tabularrobotics1K<n<10K0 likes48 downloads1y agoHugging Face16sysmlv2research /tutorials_code_and_text Tutorials Extracted Text Dataset This is the extracted text dataset of sysmlv2's official tutorials pdf. With the text explaination and code examples in each page. Useful for training LLM and teach it the basic knowledge and conceptions of sysmlv2. 1315 records, 183 pages in total. tabular1K<n<10K0 likes46 downloads2y agoHugging Face17smartloop-ai /lexic-ai-tutorial-datasettextquestion-answeringn<1K0 likes46 downloads2y agoHugging Face18ai4bharat /Spoken-Tutorialgated BhasaAnuvaad: A Speech Translation Dataset for 13 Indian Languages Overview BhasaAnuvaad, is the largest Indic-language AST dataset spanning over 44,400 hours of speech and 17M text segments for 13 of 22 scheduled Indian languages and English. This repository consists of parallel data for Speech Translation from Spoken-Tutorial youtube channel, a subset of BhasaAnuvaad. How to use The datasets library allows you to load and pre-process your… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/Spoken-Tutorial.audio100K<n<1M2 likes42 downloads2y agoHugging Face19notmahi /tutorial-ball-top-20tabular10K<n<100K0 likes40 downloads2y agoHugging Face20masakinoda /so101-tutorial-eraser-9This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 1, "total_frames": 1798, "total_tasks": 1, "total_videos": 1, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/masakinoda/so101-tutorial-eraser-9.tabularrobotics1K<n<10K0 likes39 downloads1y agoHugging Face21pravsels /Manim-Tutorials-2021_brianamedee_codetextn<1K1 likes34 downloads3y agoHugging Face22nataliaElv /setfit_tutorial Dataset Card for setfit_tutorial This dataset has been created with Argilla. As shown in the sections below, this dataset can be loaded into Argilla as explained in Load with Argilla, or used directly with the datasets library in Load with datasets. Dataset Summary This dataset contains: A dataset configuration file conforming to the Argilla dataset format named argilla.yaml. This configuration file will be used to configure the dataset when using the… See the full description on the dataset page: https://huggingface.co/datasets/nataliaElv/setfit_tutorial.text1K<n<10K0 likes33 downloads2y agoHugging Face23TutorialGuide /blended-skill-talk-fixed Compatibility Update This repository is a compatibility-fixed version of the original Blended Skill Talk dataset. The original dataset can be found at: Original Hugging Face dataset: https://huggingface.co/datasets/anezatra/blended-skill-talk This version was created to maintain compatibility with newer versions of the Hugging Face datasets library. Changes from the Original Dataset The following changes were made: Removed the unused label_candidates column.… See the full description on the dataset page: https://huggingface.co/datasets/TutorialGuide/blended-skill-talk-fixed.texttext-generation1K<n<10K0 likes33 downloads2mo agoHugging Face24nataliaElv /dolly_tutorial Dataset Card for dolly_tutorial This dataset has been created with Argilla. As shown in the sections below, this dataset can be loaded into Argilla as explained in Load with Argilla, or used directly with the datasets library in Load with datasets. Dataset Summary This dataset contains: A dataset configuration file conforming to the Argilla dataset format named argilla.cfg. This configuration file will be used to configure the dataset when using the… See the full description on the dataset page: https://huggingface.co/datasets/nataliaElv/dolly_tutorial.text10K<n<100K0 likes32 downloads3y agoHugging Face25masakinoda /so101-tutorial-eraser-13This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 2, "total_frames": 1196, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/masakinoda/so101-tutorial-eraser-13.tabularrobotics1K<n<10K0 likes32 downloads1y agoHugging Face26windchimeran /speculator-tutorial speculator-tutorial Raw vs. on-policy regenerated conversation data for training speculative-decoding drafters (EAGLE-3 / DFlash / DSpark style), with the original source data kept alongside so you can see exactly what regeneration changes and why it matters. Prompts come from UltraChat-200k. The verifier / teacher model is Qwen/Qwen3-8B. Why regenerate at all? A speculative-decoding drafter is trained to predict what the verifier would say next. If you train it… See the full description on the dataset page: https://huggingface.co/datasets/windchimeran/speculator-tutorial.tabulartext-generation1K<n<10K0 likes31 downloads2mo agoHugging Face27amtellezfernandez /robot-learning-tutorial-dataThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so100_follower", "total_episodes": 5, "total_frames": 2984, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:5" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/amtellezfernandez/robot-learning-tutorial-data.tabularrobotics1K<n<10K1 likes30 downloads10mo agoHugging Face28kimz1121 /lerobot_tutorial_2025_07_23_15_52_35This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "my_cool_robot", "total_episodes": 5, "total_frames": 186, "total_tasks": 1, "total_videos": 5, "total_chunks": 1, "chunks_size": 1000, "fps": 15, "splits": { "train": "0:5"}, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kimz1121/lerobot_tutorial_2025_07_23_15_52_35.tabularroboticsn<1K0 likes29 downloads1y agoHugging Face29masakinoda /so101-tutorial-eraser-26This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 2, "total_frames": 1192, "total_tasks": 1, "total_videos": 4, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/masakinoda/so101-tutorial-eraser-26.tabularrobotics1K<n<10K0 likes29 downloads1y agoHugging Face305hred /tutorialThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "so100", "total_episodes": 5, "total_frames": 1510, "total_tasks":1, "total_videos": 10, "total_chunks": 1, "chunks_size": 1000, "fps": 24, "splits": { "train": "0:5" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/5hred/tutorial.tabularrobotics1K<n<10K0 likes28 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.