CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01team-hatakeyama-phase2 /ndlj_tosho_1 国会図書館に収蔵される著作権切れのデータです textn<1K0 likes713 downloads2y agoHugging Face02tinkhireeva /so101_ring_tossThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 6, "total_frames": 4841, "total_tasks": 1, "total_videos": 12, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:6" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tinkhireeva/so101_ring_toss.tabularrobotics10K<n<100K0 likes142 downloads1y agoHugging Face03CodeHima /TOS_Dataset TOS_Dataset This dataset contains clauses from Terms of Service (ToS) documents with annotations indicating the fairness level of each clause. The dataset includes clauses labeled as clearly_fair, potentially_unfair, and clearly_unfair. Dataset Summary The dataset comprises clauses extracted from various ToS documents. Each clause is annotated with a fairness level, indicating whether it is clearly fair, potentially unfair, or clearly unfair. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/CodeHima/TOS_Dataset.text1K<n<10K0 likes116 downloads2y agoHugging Face04LawInformedAI /claudette_tos Dataset Card for "claudette_tos" More Information needed text1K<n<10K3 likes113 downloads3y agoHugging Face05chenghao /tos_pp_dataset A collection of Terms of Service or Privacy Policy datasets Annotated datasets CUAD Specifically, the 28 service agreements from CUAD, which are licensed under CC BY 4.0 (subset: cuad). Code import datasets from tos_datasets.proto import DocumentQA ds = datasets.load_dataset("chenghao/tos_pp_dataset", "cuad") print(DocumentQA.model_validate_json(ds["document"][0])) 100 ToS From Annotated 100 ToS, CC BY 4.0 (subset: 100_tos). Code import… See the full description on the dataset page: https://huggingface.co/datasets/chenghao/tos_pp_dataset.text10K<n<100K0 likes111 downloads2y agoHugging Face06CodeHima /TOS_DatasetV3 TOS_DatasetV3 Dataset Description TOS_DatasetV3 is a dataset designed for analyzing the unfairness of terms of service (ToS) clauses. It includes sentences from various terms of service agreements categorized into three unfairness levels: clearly_fair, potentially_unfair, and clearly_unfair. This dataset aims to aid in the development of models that can assess the fairness of legal documents. Dataset Structure The dataset consists of the following columns:… See the full description on the dataset page: https://huggingface.co/datasets/CodeHima/TOS_DatasetV3.text10K<n<100K0 likes91 downloads2y agoHugging Face07DaivdYuan /umi-dynamic-tossing-lerobot Dynamic Tossing Visualizers LeRobot Visualizer Neural Motion Visualizer Overview This dataset is a conversion from an upstream robotics dataset into LeRobot v3-compatible format. It is intended to provide reproducible access in a unified schema. Source Dataset Dataset ID: umi-dynamic-tossing Project: UMI (Core) Task: Dynamic Tossing Upstream download/source URL: https://real.stanford.edu/umi/data/dynamic_tossing/dynamic_tossing.zarr.zip Upstream… See the full description on the dataset page: https://huggingface.co/datasets/DaivdYuan/umi-dynamic-tossing-lerobot.tabular100K<n<1M0 likes67 downloads6mo agoHugging Face08Toshi-Noguchi /teleope_so101_dual_handkercheif_merged_first_iteration_ver2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "bi_so101_follower", "total_episodes": 82, "total_frames": 128010, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:82" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Toshi-Noguchi/teleope_so101_dual_handkercheif_merged_first_iteration_ver2.tabularrobotics100K<n<1M0 likes66 downloads10mo agoHugging Face09tinkhireeva /eval_so101_ring_tossThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 4, "total_frames": 4458, "total_tasks": 1, "total_videos": 8, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:4" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tinkhireeva/eval_so101_ring_toss.tabularrobotics1K<n<10K0 likes58 downloads1y agoHugging Face10Richard-Nai /HuMI-Toss Dataset Card for HuMI [Project Page] | [Paper] Dataset Summary This dataset was collected using the HuMI data collection pipeline and converted into the LeRobot format. It provides robot-free demonstrations for humanoid whole-body manipulation. Task Description: This repository features a dynamic-toss task in which the humanoid throws a toy into a target cart. The dataset consists of 104 demonstrations collected in a single environment. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/Richard-Nai/HuMI-Toss.tabularrobotics10K<n<100K0 likes53 downloads7mo agoHugging Face11toshi456 /LLaVA-CC3M-Pretrain-595K-JA Dataset Card for "LLaVA-CC3M-Pretrain-595K-JA" Dataset Details Dataset Type: Japanese LLaVA CC3M Pretrain 595K is a localized version of the original LLaVA Visual Instruct CC3M 595K dataset. This version is translated into Japanese using cyberagent/calm2-7b-chat and is aimed at serving similar purposes in the context of Japanese language. Resources for More Information: For information on the original dataset: liuhaotian/LLaVA-CC3M-Pretrain-595K License: Must comply with… See the full description on the dataset page: https://huggingface.co/datasets/toshi456/LLaVA-CC3M-Pretrain-595K-JA.textvisual-question-answering100K<n<1M11 likes50 downloads2y agoHugging Face12Toshi-Noguchi /teleope_so101_dual_handkercheif_merged_first_iterationThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "bi_so101_follower", "total_episodes": 82, "total_frames": 128010, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:82" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Toshi-Noguchi/teleope_so101_dual_handkercheif_merged_first_iteration.tabularrobotics100K<n<1M0 likes49 downloads10mo agoHugging Face13Toshi-Noguchi /teleope1020_task-dual-merged_ver2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "bi_so101_follower", "total_episodes": 4, "total_frames": 9541, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:4" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Toshi-Noguchi/teleope1020_task-dual-merged_ver2.tabularrobotics1K<n<10K0 likes48 downloads11mo agoHugging Face14Toshi-Noguchi /teleope_so101_dual_handkercheif_merged_first_iteration_v2_fixedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "bi_so101_follower", "total_episodes": 82, "total_frames": 128010, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:82" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Toshi-Noguchi/teleope_so101_dual_handkercheif_merged_first_iteration_v2_fixed.tabularrobotics100K<n<1M0 likes48 downloads10mo agoHugging Face15cfierro /c4-en-2k-tos-game-replay Fixed English C4 replay subset A subset of allenai/c4, English configuration, training split. C4 is derived from Common Crawl; see the upstream card for provenance and licensing. Sized against cfierro/tos_game_synthetic_docs, split train, using raw text tokens without special tokens or truncation. All 6,219 documents are in train, with 2,982,687 raw tokens. Whole documents are kept until the target is reached; exact duplicate texts are skipped. id is SHA-256 of the original… See the full description on the dataset page: https://huggingface.co/datasets/cfierro/c4-en-2k-tos-game-replay.tabular1K<n<10K0 likes47 downloads18d agoHugging Face16Toshi-Noguchi /teleope1020_task-dual-mergedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "bi_so101_follower", "total_episodes": 16, "total_frames": 46191, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:16" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Toshi-Noguchi/teleope1020_task-dual-merged.tabularrobotics10K<n<100K0 likes38 downloads11mo agoHugging Face17Tamazight-NLP /TOSD Dataset Card for Tamazight Open Speech Dataset This dataset provides a parsed, formatted, and ready-to-use Amazigh Voice Dataset. It contains voice recordings and corresponding text transcripts in Standard Moroccan Amazigh (ⵜⴰⵎⴰⵣⵉⵖⵜ ⵜⴰⵏⴰⵡⴰⵢⵜ ⵜⴰⵎⵓⵔⴰⴽⵓⵛⵜ) intended for training Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) models. This specific repository is published by a collaborator. You may visit the raw dataset repository which has additional dataset that hasn't… See the full description on the dataset page: https://huggingface.co/datasets/Tamazight-NLP/TOSD.audioautomatic-speech-recognition1K<n<10K3 likes34 downloads5mo agoHugging Face18Toshi-Noguchi /teleope_so101_dual_handkercheif_merged_first_iteration_v2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "bi_so101_follower", "total_episodes": 82, "total_frames": 128010, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:82" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Toshi-Noguchi/teleope_so101_dual_handkercheif_merged_first_iteration_v2.tabularrobotics100K<n<1M0 likes33 downloads10mo agoHugging Face19DaivdYuan /hub-tennis-ball-basket-toss-lerobot Tennis Ball Basket Toss Visualizers LeRobot Visualizer Neural Motion Visualizer Overview This dataset is a conversion from an upstream robotics dataset into LeRobot v3-compatible format. It is intended to provide reproducible access in a unified schema. Source Dataset Dataset ID: hub-tennis-ball-basket-toss Project: unknown Task: Tennis Ball Basket Toss Upstream download/source URL: https://real.stanford.edu/umi-on-legs/tossing.zarr.zip Upstream… See the full description on the dataset page: https://huggingface.co/datasets/DaivdYuan/hub-tennis-ball-basket-toss-lerobot.tabular100K<n<1M0 likes32 downloads6mo agoHugging Face20qanta-challenge /acf-co24-tossupstextn<1K0 likes26 downloads1y agoHugging Face21tostideluxekaas /toxic-dpo-v0.2-dutch Toxic DPO v0.2 - Dutch Translation Dataset Description This is a direct machine-translated Dutch version of the original datasetunalignment/toxic-dpo-v0.2. Translation method:English → Dutch using the Helsinki-NLP/opus-mt-en-nl model from the MarianMTModel translations.No manual edits, additions or filtering were applied besides the automated translation. Data set is checked on NULL values and duplicates. All fields (prompt, chosen, rejected) were translated… See the full description on the dataset page: https://huggingface.co/datasets/tostideluxekaas/toxic-dpo-v0.2-dutch.textn<1K3 likes25 downloads7mo agoHugging Face22danodev /toss-tennis-ball-merged2-v2.1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 100, "total_frames": 41127, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:100" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/danodev/toss-tennis-ball-merged2-v2.1.tabularrobotics10K<n<100K0 likes24 downloads10mo agoHugging Face23mgor /acf-co24-tossupstext1K<n<10K1 likes23 downloads1y agoHugging Face24iarbel /unfair_tos_fewshot_eval Dataset Card for "unfair_tos_fewshot_eval" More Information needed text1K<n<10K0 likes21 downloads2y agoHugging Face25CodeHima /TOS_DatasetV2text10K<n<100K0 likes21 downloads2y agoHugging Face26tinkhireeva /eval_so101_ring_toss_actThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 2, "total_frames": 1557, "total_tasks": 1, "total_videos": 4, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tinkhireeva/eval_so101_ring_toss_act.tabularrobotics1K<n<10K0 likes21 downloads1y agoHugging Face27tosiyama /beans Dataset Card for Beans Dataset Summary Beans leaf dataset with images of diseased and health leaves. Supported Tasks and Leaderboards image-classification: Based on a leaf image, the goal of this task is to predict the disease type (Angular Leaf Spot and Bean Rust), if any. Languages English Dataset Structure Data Instances A sample from the training set is provided below: { 'image_file_path':… See the full description on the dataset page: https://huggingface.co/datasets/tosiyama/beans.imageimage-classification1K<n<10K0 likes20 downloads9mo agoHugging Face28prasannadhungana8848 /TOS_sentence_embedded_all_minilm_l6_v2text10K<n<100K0 likes19 downloads2y agoHugging Face29toshirobot /dart-eraser-placementThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 35, "total_frames": 14636, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:35" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/toshirobot/dart-eraser-placement.tabularrobotics10K<n<100K0 likes13 downloads6mo agoHugging Face30tosoham /recipe-datasetimagen<1K0 likes12 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.