CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01twangodev /librivox-mirror LibriVox Mirror Fast, structured, continuously updated LibriVox audio mirror. Current snapshot Metric Value Published books 21,734 Published sections 493,396 Audio hours 132,613.0 Audio languages 86 Quarantined books 603 Last updated (UTC) 2026-09-24T13:43:06.550876Z Audio by language Language Hours English 131,663.6 German 417.0 Spanish 160.9 French 103.8 Portuguese 37.4 Polish 34.1 Dutch 25.8… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/librivox-mirror.audioautomatic-speech-recognition100K<n<1M0 likes23k downloads16h agoHugging Face020xKai /mirador-offloadtabular100M<n<1B0 likes6.3k downloads25d agoHugging Face03leeaandrob /mirror-eduagarcia__CrawlPT_dedup CrawlPT (deduplicated) CrawlPT is a generic Portuguese corpus extracted from various web pages. This version is deduplicated using MinHash algorithm and Locality Sensitive Hashing, following the approach of Lee et al. (2022). The raw version is also available here. Dataset Details Dataset is composed by three corpora: brWaC, C100-PT, OSCAR-2301. brWaC: a web corpus for Brazilian Portuguese from 120,000 different websites. C100-PT: Portuguese subset from CC-100.… See the full description on the dataset page: https://huggingface.co/datasets/leeaandrob/mirror-eduagarcia__CrawlPT_dedup.tabulartext-generation100M<n<1B0 likes1.8k downloads3mo agoHugging Face04nvidia /miracl-vision MIRACL-VISION MIRACL-VISION is a multilingual visual retrieval dataset for 18 different languages. It is an extension of MIRACL, a popular text-only multilingual retrieval dataset. The dataset contains user questions, images of Wikipedia articles and annotations, which article can answer a user question. There are 7898 questions and 338734 images. More details can be found in the paper MIRACL-VISION: A Large, multilingual, visual document retrieval benchmark. This dataset is ready… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/miracl-vision.image100K<n<1M13 likes1.3k downloads1y agoHugging Face05brhkim /education_data_portal_mirror_2026q3 Education Data Portal — Parquet Mirror (2026Q3 · Portal v0.26.1) A complete mirror of the Urban Institute Education Data Portal datasets version 0.26.1, collected August 6, 2026, and converted from CSV to Apache Parquet format for efficient analytical use. Please note that the maintainers of this Huggingface Dataset have no affiliation with the Urban Institute or the Education Data Portal team. Huge appreciation for all they do -- if you use this mirror, please make sure to… See the full description on the dataset page: https://huggingface.co/datasets/brhkim/education_data_portal_mirror_2026q3.tabular1B<n<10B0 likes990 downloads2mo agoHugging Face06aipracticecafe-mirror /curated-danbooru-2026-256px-flux2-vaetabular100K<n<1M0 likes820 downloads29d agoHugging Face07TencentARC /MiraData MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions Xuan Ju1*, Yiming Gao1*, Zhaoyang Zhang1*#, Ziyang Yuan1, Xintao Wang1, Ailing Zeng, Yu Xiong, Qiang Xu, Ying Shan1 1ARC Lab, Tencent PCG 2The Chinese University of Hong Kong *Equal Contribution #Project Lead Introduction Video datasets play a crucial role in video generation such as Sora. However, existing text-video datasets often fall short when it comes to handling long video… See the full description on the dataset page: https://huggingface.co/datasets/TencentARC/MiraData.tabularimage-to-video100K<n<1M42 likes452 downloads2y agoHugging Face08mulab-mir /muchomusic MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models MuChoMusic is a benchmark designed to evaluate music understanding in multimodal language models focused on audio. It includes 1,187 multiple-choice questions validated by human annotators, based on 644 music tracks from two publicly available music datasets. These questions cover a wide variety of genres and assess knowledge and reasoning across several musical concepts and their cultural and functional… See the full description on the dataset page: https://huggingface.co/datasets/mulab-mir/muchomusic.tabular1K<n<10K8 likes414 downloads2y agoHugging Face09Mireu-Lab /NSL-KDD NSL-KDD The data set is a data set that converts the arff File provided by the link into CSV and results. The data set is personally stored by converting data to float64. If you want to obtain additional original files, they are organized in the Original Directory in the repo. Labels The label of the data set is as follows. # Column Non-Null Count Dtype 0 duration 151165 non-null int64 1 protocol_type 151165 non-null object 2 service 151165 non-null… See the full description on the dataset page: https://huggingface.co/datasets/Mireu-Lab/NSL-KDD.tabular100K<n<1M6 likes381 downloads2y agoHugging Face10ZhuofengLi /MiroVerse-v0.1tabular100K<n<1M0 likes325 downloads8mo agoHugging Face11MirandaAbhilash /vqav2-full-metadatatabular100K<n<1M0 likes322 downloads6mo agoHugging Face12Mireu-Lab /UNSW-NB15 UNSW-NB15 This data is provided through the Train, Test CSV file provided by UNSW-NB15. link Labels The label of the data set is as follows. # Column Non-Null Count Dtype 0 id 82332 non-null int64 1 dur 82332 non-null float64 2 proto 82332 non-null object 3 service 82332 non-null object 4 state 82332 non-null object 5 spkts 82332 non-null int64 6 dpkts 82332 non-null int64 7 sbytes 82332 non-null int64 8 dbytes 82332 non-null int64 9 rate… See the full description on the dataset page: https://huggingface.co/datasets/Mireu-Lab/UNSW-NB15.tabular100K<n<1M1 likes283 downloads2y agoHugging Face13leeaandrob /mirror-eduagarcia__LegalPT_dedup LegalPT (deduplicated) LegalPT aggregates the maximum amount of publicly available legal data in Portuguese, drawing from varied sources including legislation, jurisprudence, legal articles, and government documents. This version is deduplicated using MinHash algorithm and Locality Sensitive Hashing, following the approach of Lee et al. (2022). The raw version is also available here. Dataset Details Dataset is composed by six corpora: Ulysses-Tesemõ… See the full description on the dataset page: https://huggingface.co/datasets/leeaandrob/mirror-eduagarcia__LegalPT_dedup.tabular10M<n<100M0 likes201 downloads3mo agoHugging Face14lerobot-data-collection /level12_rac_2_2026-02-07_and_mirThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "openarms_follower", "total_episodes": 10770, "total_frames": 26748966, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:10770" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/level12_rac_2_2026-02-07_and_mir.tabularrobotics10M<n<100M1 likes169 downloads8mo agoHugging Face15brhkim /education_data_portal_mirror ⚠️ Frozen Vintage — Education Data Portal Parquet Mirror (v0.24.0, February 2026) This repository is frozen and will receive no further updates. It is preserved permanently as a snapshot for reproducibility. For current data, use the successor mirror: brhkim/education_data_portal_mirror_2026q3 (Education Data Portal v0.26.1, collected August 2026). If you are reproducing an analysis that originally used this mirror, you are in the right place — keep your scripts pointed here.… See the full description on the dataset page: https://huggingface.co/datasets/brhkim/education_data_portal_mirror.tabular1B<n<10B0 likes167 downloads2mo agoHugging Face16Kiva12138 /mirth_lerobot MIRTH Dataset Multi-camera real-world manipulation demonstrations for history-aware Vision-Language-Action agents The MIRTH dataset is a real-world robot manipulation dataset collected on a physical LeRobot platform. It contains synchronized main-camera and wrist-camera observations, robot proprioception, language instructions, and expert action trajectories for training and evaluating Vision-Language-Action (VLA) agents. This release provides the same… See the full description on the dataset page: https://huggingface.co/datasets/Kiva12138/mirth_lerobot.tabular100K<n<1M0 likes163 downloads3mo agoHugging Face17queyuecanyang /MIRACLE MIRACLE MIRACLE is a multimodal benchmark dataset with image-based questions and model evaluation results. This repository contains the test subset of the MIRACLE benchmark. The benchmark config provides the benchmark test split, and the model_results config provides test-set evaluation outputs for the included models. Dataset Structure The repository is organized as follows: data/ test.parquet # Hugging Face loadable benchmark test split test.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/queyuecanyang/MIRACLE.imagevisual-question-answering1K<n<10K4 likes152 downloads3mo agoHugging Face18OpenVoiceOS /ovos-wake-word-bench-picovoice-smart-mirror OVOS wake_word bench — picovoice-smart-mirror Per-clip detection decisions predictions of the registered OVOS Plugin Arena wake_word fighters over Picovoice/wake-word-benchmark. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo; the… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-smart-mirror.tabular1K<n<10K0 likes144 downloads16d agoHugging Face19LevenKoko /MIRAGE-CanaryDocs MIRAGE CanaryDocs MIRAGE CanaryDocs is an English synthetic enterprise-document dataset for structured privacy-unit, canary, and ordered multi-chunk evaluation. It is the companion dataset for the EMNLP 2026 paper When Metadata Remembers: Ordered Provenance Enables Document-Level Embedding Inversion. Project documentation and schemas are also available in the MIRAGE GitHub repository. Dataset summary The dataset contains complete synthetic documents, ordered token… See the full description on the dataset page: https://huggingface.co/datasets/LevenKoko/MIRAGE-CanaryDocs.tabulartoken-classification1M<n<10M0 likes132 downloads26d agoHugging Face20villekuosmanen /agilex_push_mirror_surfaceThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "arx5_bimanual", "total_episodes": 20, "total_frames": 6503, "total_tasks": 1, "total_videos": 60, "total_chunks": 1, "chunks_size": 1000, "fps": 25, "splits": { "train": "0:20" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/agilex_push_mirror_surface.tabularrobotics1K<n<10K0 likes130 downloads7mo agoHugging Face21lurcelay /mir2023tabular10K<n<100K0 likes123 downloads2y agoHugging Face22stabletoolbench /MirrorAPI-Training MirrorAPI training dataset This dataset contains the training data for MirrorAPI and MirrorAPI-Cache: train_sft.json, train_cot.json, train_augment.json: The training data for MirrorAPI . train_cache.json: The training data for MirrorAPI-Cache. tabular100K<n<1M1 likes112 downloads2y agoHugging Face23villekuosmanen /agilex_cover_mirror_jshThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "arx5_bimanual", "total_episodes": 20, "total_frames": 7153, "total_tasks": 1, "total_videos": 60, "total_chunks": 1, "chunks_size": 1000, "fps": 25, "splits": { "train": "0:20" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/agilex_cover_mirror_jsh.tabularrobotics1K<n<10K0 likes110 downloads7mo agoHugging Face24rlhf-and-friends /proofwriter-mirror ProofWriter (The Mirror) A typed, content-faithful mirror of AI2's ProofWriter dataset (release V2020.12.3), derived from proofwriter-source. The JSON encoding is cleaned up: the id-keyed dicts (triple1, Q3, …) become lists of structs that keep their id, every atom representation is parsed into a typed {subject, relation, object, polarity} triple, and the closed enums (answer, strategy) are typed. The content stays faithful — nothing renamed, no rows dropped — and the recursive… See the full description on the dataset page: https://huggingface.co/datasets/rlhf-and-friends/proofwriter-mirror.tabular100K<n<1M0 likes98 downloads25d agoHugging Face25alphabot2 /03-09-Left-ReGrasp_mirrored_right_armThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "aibot2", "total_episodes": 57, "total_frames": 4552, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 10, "splits": { "train": "0:57" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/alphabot2/03-09-Left-ReGrasp_mirrored_right_arm.tabularrobotics1K<n<10K0 likes96 downloads21d agoHugging Face26driodnexus /backln-guest-post-quality-public-mirror Backln Guest Post Quality Public Mirror Public-safe mirror for validating Hugging Face Dataset Viewer indexing and release gates. This dataset is not the private training corpus. Full text, titles, and snippets are removed by default. The mirror keeps labels, coarse metadata, feature buckets, and hash prefixes so the public Hub can verify schema and distribution without exposing customer content. Schema label: one of published, manual_review, rejected. source: coarse… See the full description on the dataset page: https://huggingface.co/datasets/driodnexus/backln-guest-post-quality-public-mirror.tabulartext-classificationn<1K0 likes92 downloads4mo agoHugging Face27MirkoSchiavone /mawi-https-flows-2025tabular1M<n<10M0 likes91 downloads1y agoHugging Face28rainbowrobotics /flowers_sorting_mirroredThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "right_arm_0", "right_arm_1", "right_arm_2", "right_arm_3", "right_arm_4", "right_arm_5", "right_arm_6", "left_arm_0"… See the full description on the dataset page: https://huggingface.co/datasets/rainbowrobotics/flowers_sorting_mirrored.tabularrobotics100K<n<1M0 likes90 downloads23d agoHugging Face29miria0 /EduFeedback EduFeedback Alternating dataset example: a single curated multi-turn conversation yields a complete (prompt, chosen, rejected) triplet on its own — the direct early response becomes chosen and a later, less-direct response becomes rejected. Both sides come from the same real dialog, so no synthetic LLM generation is needed to fill in the rejected side. EduFeedback is a synthetically generated, multi-turn conversational preference dataset in an educational tutoring setting… See the full description on the dataset page: https://huggingface.co/datasets/miria0/EduFeedback.tabulartext-generation100K<n<1M0 likes88 downloads4mo agoHugging Face30Screener2 /V4_pnp_combined_mirroredThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/Screener2/V4_pnp_combined_mirrored.tabularrobotics10K<n<100K0 likes86 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.