CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01twangodev /librivox-mirror LibriVox Mirror Fast, structured, continuously updated LibriVox audio mirror. Current snapshot Metric Value Published books 21,725 Published sections 493,206 Audio hours 132,555.1 Audio languages 86 Quarantined books 609 Last updated (UTC) 2026-09-22T13:24:09.765701Z Audio by language Language Hours English 131,605.7 German 417.0 Spanish 160.9 French 103.8 Portuguese 37.4 Polish 34.1 Dutch 25.8… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/librivox-mirror.audioautomatic-speech-recognition100K<n<1M0 likes23k downloads10h agoHugging Face020xKai /mirador-offloadtabular100M<n<1B0 likes7.5k downloads23d agoHugging Face03leeaandrob /mirror-eduagarcia__CrawlPT_dedup CrawlPT (deduplicated) CrawlPT is a generic Portuguese corpus extracted from various web pages. This version is deduplicated using MinHash algorithm and Locality Sensitive Hashing, following the approach of Lee et al. (2022). The raw version is also available here. Dataset Details Dataset is composed by three corpora: brWaC, C100-PT, OSCAR-2301. brWaC: a web corpus for Brazilian Portuguese from 120,000 different websites. C100-PT: Portuguese subset from CC-100.… See the full description on the dataset page: https://huggingface.co/datasets/leeaandrob/mirror-eduagarcia__CrawlPT_dedup.tabulartext-generation100M<n<1B0 likes1.4k downloads3mo agoHugging Face04nvidia /miracl-vision MIRACL-VISION MIRACL-VISION is a multilingual visual retrieval dataset for 18 different languages. It is an extension of MIRACL, a popular text-only multilingual retrieval dataset. The dataset contains user questions, images of Wikipedia articles and annotations, which article can answer a user question. There are 7898 questions and 338734 images. More details can be found in the paper MIRACL-VISION: A Large, multilingual, visual document retrieval benchmark. This dataset is ready… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/miracl-vision.image100K<n<1M13 likes1.2k downloads1y agoHugging Face05brhkim /education_data_portal_mirror_2026q3 Education Data Portal — Parquet Mirror (2026Q3 · Portal v0.26.1) A complete mirror of the Urban Institute Education Data Portal datasets version 0.26.1, collected August 6, 2026, and converted from CSV to Apache Parquet format for efficient analytical use. Please note that the maintainers of this Huggingface Dataset have no affiliation with the Urban Institute or the Education Data Portal team. Huge appreciation for all they do -- if you use this mirror, please make sure to… See the full description on the dataset page: https://huggingface.co/datasets/brhkim/education_data_portal_mirror_2026q3.tabular1B<n<10B0 likes906 downloads2mo agoHugging Face06aipracticecafe-mirror /curated-danbooru-2026-256px-flux2-vaetabular100K<n<1M0 likes820 downloads26d agoHugging Face07TencentARC /MiraData MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions Xuan Ju1*, Yiming Gao1*, Zhaoyang Zhang1*#, Ziyang Yuan1, Xintao Wang1, Ailing Zeng, Yu Xiong, Qiang Xu, Ying Shan1 1ARC Lab, Tencent PCG 2The Chinese University of Hong Kong *Equal Contribution #Project Lead Introduction Video datasets play a crucial role in video generation such as Sora. However, existing text-video datasets often fall short when it comes to handling long video… See the full description on the dataset page: https://huggingface.co/datasets/TencentARC/MiraData.tabularimage-to-video100K<n<1M42 likes464 downloads2y agoHugging Face08Mireu-Lab /NSL-KDD NSL-KDD The data set is a data set that converts the arff File provided by the link into CSV and results. The data set is personally stored by converting data to float64. If you want to obtain additional original files, they are organized in the Original Directory in the repo. Labels The label of the data set is as follows. # Column Non-Null Count Dtype 0 duration 151165 non-null int64 1 protocol_type 151165 non-null object 2 service 151165 non-null… See the full description on the dataset page: https://huggingface.co/datasets/Mireu-Lab/NSL-KDD.tabular100K<n<1M6 likes397 downloads2y agoHugging Face09mulab-mir /muchomusic MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models MuChoMusic is a benchmark designed to evaluate music understanding in multimodal language models focused on audio. It includes 1,187 multiple-choice questions validated by human annotators, based on 644 music tracks from two publicly available music datasets. These questions cover a wide variety of genres and assess knowledge and reasoning across several musical concepts and their cultural and functional… See the full description on the dataset page: https://huggingface.co/datasets/mulab-mir/muchomusic.tabular1K<n<10K8 likes390 downloads2y agoHugging Face10MirandaAbhilash /vqav2-full-metadatatabular100K<n<1M0 likes369 downloads6mo agoHugging Face11Mireu-Lab /UNSW-NB15 UNSW-NB15 This data is provided through the Train, Test CSV file provided by UNSW-NB15. link Labels The label of the data set is as follows. # Column Non-Null Count Dtype 0 id 82332 non-null int64 1 dur 82332 non-null float64 2 proto 82332 non-null object 3 service 82332 non-null object 4 state 82332 non-null object 5 spkts 82332 non-null int64 6 dpkts 82332 non-null int64 7 sbytes 82332 non-null int64 8 dbytes 82332 non-null int64 9 rate… See the full description on the dataset page: https://huggingface.co/datasets/Mireu-Lab/UNSW-NB15.tabular100K<n<1M1 likes287 downloads2y agoHugging Face12ZhuofengLi /MiroVerse-v0.1tabular100K<n<1M0 likes211 downloads8mo agoHugging Face13leeaandrob /mirror-eduagarcia__LegalPT_dedup LegalPT (deduplicated) LegalPT aggregates the maximum amount of publicly available legal data in Portuguese, drawing from varied sources including legislation, jurisprudence, legal articles, and government documents. This version is deduplicated using MinHash algorithm and Locality Sensitive Hashing, following the approach of Lee et al. (2022). The raw version is also available here. Dataset Details Dataset is composed by six corpora: Ulysses-Tesemõ… See the full description on the dataset page: https://huggingface.co/datasets/leeaandrob/mirror-eduagarcia__LegalPT_dedup.tabular10M<n<100M0 likes162 downloads3mo agoHugging Face14Kiva12138 /mirth_lerobot MIRTH Dataset Multi-camera real-world manipulation demonstrations for history-aware Vision-Language-Action agents The MIRTH dataset is a real-world robot manipulation dataset collected on a physical LeRobot platform. It contains synchronized main-camera and wrist-camera observations, robot proprioception, language instructions, and expert action trajectories for training and evaluating Vision-Language-Action (VLA) agents. This release provides the same… See the full description on the dataset page: https://huggingface.co/datasets/Kiva12138/mirth_lerobot.tabular100K<n<1M0 likes150 downloads3mo agoHugging Face15queyuecanyang /MIRACLE MIRACLE MIRACLE is a multimodal benchmark dataset with image-based questions and model evaluation results. This repository contains the test subset of the MIRACLE benchmark. The benchmark config provides the benchmark test split, and the model_results config provides test-set evaluation outputs for the included models. Dataset Structure The repository is organized as follows: data/ test.parquet # Hugging Face loadable benchmark test split test.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/queyuecanyang/MIRACLE.imagevisual-question-answering1K<n<10K4 likes146 downloads3mo agoHugging Face16OpenVoiceOS /ovos-wake-word-bench-picovoice-smart-mirror OVOS wake_word bench — picovoice-smart-mirror Per-clip detection decisions predictions of the registered OVOS Plugin Arena wake_word fighters over Picovoice/wake-word-benchmark. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo; the… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-smart-mirror.tabular1K<n<10K0 likes142 downloads14d agoHugging Face17LevenKoko /MIRAGE-CanaryDocs MIRAGE CanaryDocs MIRAGE CanaryDocs is an English synthetic enterprise-document dataset for structured privacy-unit, canary, and ordered multi-chunk evaluation. It is the companion dataset for the EMNLP 2026 paper When Metadata Remembers: Ordered Provenance Enables Document-Level Embedding Inversion. Project documentation and schemas are also available in the MIRAGE GitHub repository. Dataset summary The dataset contains complete synthetic documents, ordered token… See the full description on the dataset page: https://huggingface.co/datasets/LevenKoko/MIRAGE-CanaryDocs.tabulartoken-classification1M<n<10M0 likes131 downloads24d agoHugging Face18villekuosmanen /agilex_push_mirror_surfaceThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "arx5_bimanual", "total_episodes": 20, "total_frames": 6503, "total_tasks": 1, "total_videos": 60, "total_chunks": 1, "chunks_size": 1000, "fps": 25, "splits": { "train": "0:20" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/agilex_push_mirror_surface.tabularrobotics1K<n<10K0 likes126 downloads7mo agoHugging Face19lerobot-data-collection /level12_rac_2_2026-02-07_and_mirThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "openarms_follower", "total_episodes": 10770, "total_frames": 26748966, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:10770" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/level12_rac_2_2026-02-07_and_mir.tabularrobotics10M<n<100M1 likes118 downloads8mo agoHugging Face20driodnexus /backln-guest-post-quality-public-mirror Backln Guest Post Quality Public Mirror Public-safe mirror for validating Hugging Face Dataset Viewer indexing and release gates. This dataset is not the private training corpus. Full text, titles, and snippets are removed by default. The mirror keeps labels, coarse metadata, feature buckets, and hash prefixes so the public Hub can verify schema and distribution without exposing customer content. Schema label: one of published, manual_review, rejected. source: coarse… See the full description on the dataset page: https://huggingface.co/datasets/driodnexus/backln-guest-post-quality-public-mirror.tabulartext-classificationn<1K0 likes110 downloads4mo agoHugging Face21lurcelay /mir2023tabular10K<n<100K0 likes109 downloads2y agoHugging Face22villekuosmanen /agilex_cover_mirror_jshThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "arx5_bimanual", "total_episodes": 20, "total_frames": 7153, "total_tasks": 1, "total_videos": 60, "total_chunks": 1, "chunks_size": 1000, "fps": 25, "splits": { "train": "0:20" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/agilex_cover_mirror_jsh.tabularrobotics1K<n<10K0 likes105 downloads7mo agoHugging Face23brhkim /education_data_portal_mirror ⚠️ Frozen Vintage — Education Data Portal Parquet Mirror (v0.24.0, February 2026) This repository is frozen and will receive no further updates. It is preserved permanently as a snapshot for reproducibility. For current data, use the successor mirror: brhkim/education_data_portal_mirror_2026q3 (Education Data Portal v0.26.1, collected August 2026). If you are reproducing an analysis that originally used this mirror, you are in the right place — keep your scripts pointed here.… See the full description on the dataset page: https://huggingface.co/datasets/brhkim/education_data_portal_mirror.tabular1B<n<10B0 likes105 downloads2mo agoHugging Face24stabletoolbench /MirrorAPI-Training MirrorAPI training dataset This dataset contains the training data for MirrorAPI and MirrorAPI-Cache: train_sft.json, train_cot.json, train_augment.json: The training data for MirrorAPI . train_cache.json: The training data for MirrorAPI-Cache. tabular100K<n<1M1 likes104 downloads2y agoHugging Face25arqa39 /proofwriter-mirror ProofWriter (The Mirror) A typed, content-faithful mirror of AI2's ProofWriter dataset (release V2020.12.3), derived from proofwriter-source. The JSON encoding is cleaned up: the id-keyed dicts (triple1, Q3, …) become lists of structs that keep their id, every atom representation is parsed into a typed {subject, relation, object, polarity} triple, and the closed enums (answer, strategy) are typed. The content stays faithful — nothing renamed, no rows dropped — and the recursive… See the full description on the dataset page: https://huggingface.co/datasets/arqa39/proofwriter-mirror.tabular100K<n<1M0 likes99 downloads29d agoHugging Face26rlhf-and-friends /proofwriter-mirror ProofWriter (The Mirror) A typed, content-faithful mirror of AI2's ProofWriter dataset (release V2020.12.3), derived from proofwriter-source. The JSON encoding is cleaned up: the id-keyed dicts (triple1, Q3, …) become lists of structs that keep their id, every atom representation is parsed into a typed {subject, relation, object, polarity} triple, and the closed enums (answer, strategy) are typed. The content stays faithful — nothing renamed, no rows dropped — and the recursive… See the full description on the dataset page: https://huggingface.co/datasets/rlhf-and-friends/proofwriter-mirror.tabular100K<n<1M0 likes95 downloads22d agoHugging Face27alphabot2 /03-09-Left-ReGrasp_mirrored_right_armThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "aibot2", "total_episodes": 57, "total_frames": 4552, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 10, "splits": { "train": "0:57" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/alphabot2/03-09-Left-ReGrasp_mirrored_right_arm.tabularrobotics1K<n<10K0 likes94 downloads19d agoHugging Face28Screener2 /V4_pnp_combined_mirroredThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/Screener2/V4_pnp_combined_mirrored.tabularrobotics10K<n<100K0 likes93 downloads2mo agoHugging Face29mirzaei2114 /stackoverflowVQA Dataset Card for "stackoverflowVQA" More Information needed tabularvisual-question-answering1M<n<10M5 likes88 downloads3y agoHugging Face30MirkoSchiavone /mawi-https-flows-2025tabular1M<n<10M0 likes88 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.