CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01twangodev /librivox-mirror LibriVox Mirror Fast, structured, continuously updated LibriVox audio mirror. Current snapshot Metric Value Published books 21,725 Published sections 493,206 Audio hours 132,555.1 Audio languages 86 Quarantined books 609 Last updated (UTC) 2026-09-22T13:24:09.765701Z Audio by language Language Hours English 131,605.7 German 417.0 Spanish 160.9 French 103.8 Portuguese 37.4 Polish 34.1 Dutch 25.8… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/librivox-mirror.audioautomatic-speech-recognition100K<n<1M0 likes23k downloads19h agoHugging Face02DannHiroaki /China-Building-Footprints-CMAB-Mirror Origin Data @misc{Zhang2025CMAB, author = {Zhang, Yecheng and Zhao, Huimin and Long, Ying}, title = {{CMAB-The World's First National-Scale Multi-Attribute Building Dataset}}, year = {2025}, month = apr, publisher = {figshare}, doi = {10.6084/m9.figshare.27992417}, url = {https://doi.org/10.6084/m9.figshare.27992417}, howpublished = {dataset} } Paper @article{Zhang2025SciData, author = {Zhang, Y. and… See the full description on the dataset page: https://huggingface.co/datasets/DannHiroaki/China-Building-Footprints-CMAB-Mirror.geospatialn<1K0 likes13k downloads8mo agoHugging Face03mteb /MIRACLRetrievalHardNegatives MIRACLRetrievalHardNegatives An MTEB dataset Massive Text Embedding Benchmark MIRACL (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual retrieval dataset that focuses on search across 18 different languages. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct. Task category t2t Domains Encyclopaedic, Written Reference http://miracl.ai/… See the full description on the dataset page: https://huggingface.co/datasets/mteb/MIRACLRetrievalHardNegatives.texttext-retrieval1M<n<10M3 likes9.4k downloads7mo agoHugging Face04Samuelsantos777 /psg-audio-v3-unofficial-mirror PSG-Audio v3 — Unofficial Complete Mirror Unofficial complete mirror of the publicly released PSG-Audio Version 3 dataset. This repository preserves the original files without modification and provides a reliable, high-speed mirror through the Hugging Face Hub for the research community. Overview PSG-Audio v3 is one of the largest publicly available multimodal sleep datasets, combining overnight clinical polysomnography (PSG) with synchronized environmental… See the full description on the dataset page: https://huggingface.co/datasets/Samuelsantos777/psg-audio-v3-unofficial-mirror.textaudio-classificationn<1K1 likes7.7k downloads3mo agoHugging Face050xKai /mirador-offloadtabular100M<n<1B0 likes7.5k downloads23d agoHugging Face06mirobody /ESL-Bench ESL-bench ESL-bench (Event-driven Synthetic Longitudinal Benchmark) is a virtual health user dataset for evaluating AI health assistants. Each virtual user contains a complete health profile, event timeline, clinical exam data, and knowledge-graph-grounded evaluation queries, designed for use with the Mirobody-Eval framework. ⚠️ Research use only. Outputs are synthetic and intended for benchmarking AI agents. They should not be used for diagnosis or treatment decisions.… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/ESL-Bench.textquestion-answering1K<n<10K18 likes6.7k downloads16d agoHugging Face07sentence-transformers /miracl Dataset Card for MIRACL This is a reformatting of the MIRACL dataset used to train the BGE-M3 model. See the full BGE-M3 dataset in Shitao/bge-m3-data. Dataset Subsets ...-triplet subset Columns: "anchor", "positive", "negative" Column types: str, str, str Examples:{ 'anchor': '月球到地球的距离是多少?', 'positive': '月球距離\n月球距離 (LD) 是天文學上從地球到月球的距離,從地球到月球的平均距離是384,401公里 (238,856英里)。因為月球在橢圓軌道上運動,實際的距離隨時都在變化著。', 'negative':… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/miracl.textfeature-extraction1M<n<10M3 likes6.2k downloads2y agoHugging Face08mirpri /LPNSR LPNSR Dataset This repository contains the evaluation datasets and testing data associated with the paper LPNSR: Optimal Noise-Guided Diffusion Image Super-Resolution Via Learnable Noise Prediction. Project Links Paper: arXiv:2603.21045 GitHub Repository: Faze-Hsw/LPNSR Dataset Description This dataset collection is used to evaluate image super-resolution models on both synthetic and complex real-world degradations. It contains pairs of Low-Quality (LQ) and… See the full description on the dataset page: https://huggingface.co/datasets/mirpri/LPNSR.imageimage-to-image1K<n<10K0 likes6.2k downloads5mo agoHugging Face09miracl /miracl-corpus Dataset Card for MIRACL Corpus MIRACL 🌍🙌🌏 (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual retrieval dataset that focuses on search across 18 different languages, which collectively encompass over three billion native speakers around the world. This dataset contains the collection data of the 16 "known languages". The remaining 2 "surprise languages" will not be released until later. The corpus for each language is prepared from a Wikipedia… See the full description on the dataset page: https://huggingface.co/datasets/miracl/miracl-corpus.texttext-retrieval10M<n<100M54 likes4.5k downloads4y agoHugging Face10mirobody /MedHall-Bench MedHall-Bench MedHall-Bench is a field-grounded hallucination detection benchmark for medical AI assistants. It decomposes each clinical response into verifiable structured fields (dose value, unit, reference range, ICD/LOINC code, entity relation, ...) and evaluates AI outputs via per-field programmatic matching in addition to sentence-level LLM-as-Judge. Designed for use with the HolyEval framework. ⚠️ Research use only. Content is for benchmarking AI agents and should not be… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/MedHall-Bench.textquestion-answeringn<1K6 likes4.5k downloads1mo agoHugging Face11mirobody /MedHarm-Bench MedHarm-Bench MedHarm-Bench is a red-team compliance benchmark for health-management AI assistants. It uses natural-sounding patient questions that bait the assistant into crossing medical safety boundaries, then scores each response against compliance red lines. Designed for use with the HolyEval framework. ⚠️ Research use only. Questions are designed to elicit unsafe behavior for benchmarking purposes and should not be used for diagnosis or treatment decisions.… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/MedHarm-Bench.textquestion-answeringn<1K2 likes4.3k downloads1mo agoHugging Face12liuhangbiao /China-Building-Footprints-CMAB-Mirror Origin Data @misc{Zhang2025CMAB, author = {Zhang, Yecheng and Zhao, Huimin and Long, Ying}, title = {{CMAB-The World's First National-Scale Multi-Attribute Building Dataset}}, year = {2025}, month = apr, publisher = {figshare}, doi = {10.6084/m9.figshare.27992417}, url = {https://doi.org/10.6084/m9.figshare.27992417}, howpublished = {dataset} } Paper @article{Zhang2025SciData, author = {Zhang, Y. and… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/China-Building-Footprints-CMAB-Mirror.geospatialn<1K0 likes3.8k downloads6mo agoHugging Face13mteb /MIRACLRetrieval MIRACLRetrieval An MTEB dataset Massive Text Embedding Benchmark MIRACL (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual retrieval dataset that focuses on search across 18 different languages. Task category t2t Domains Encyclopaedic, Written Reference http://miracl.ai/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/MIRACLRetrieval.texttext-retrieval100M<n<1B9 likes2.9k downloads1y agoHugging Face14rtarun1 /mirage18k [IROS 2026] Mirage 18k: Dataset for Glass Segmentation & Depth Estimation Mirage 18k is a novel, multi-task dataset comprising 18,353 manually annotated images across 38 unique indoor scenes, designed specifically for joint glass segmentation and glass-aware monocular depth estimation in robotics. It contains diverse real-world glass structures (indoor panes, frosted doors, windows, clear doors) with severe background clutter, saliency, and dynamic obstacles. Model Checkpoint:… See the full description on the dataset page: https://huggingface.co/datasets/rtarun1/mirage18k.imageimage-segmentation10K<n<100K1 likes2.8k downloads2mo agoHugging Face15mirav /anime-syntheticsMostly unfiltered anime-style images generated by various text to image models, collected from various sources (some were submitted for inclusion by their creators). Includes a subset of p1atdev/niji-v5, albeit captioned differently than the source. Contains 2224 image & caption pairs. As it is unfiltered, some adult content may be included. Captions may not be completely accurate. If you wish to submit content, do it as a pull request. imagetext-to-imagen<1K6 likes2.1k downloads3y agoHugging Face16RSamoed /MIRACLRetrieval MIRACLRetrieval An MTEB dataset Massive Text Embedding Benchmark MIRACL (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual retrieval dataset that focuses on search across 18 different languages. Task category t2t Domains Encyclopaedic, Written Reference http://miracl.ai/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/RSamoed/MIRACLRetrieval.texttext-retrieval100M<n<1B0 likes1.8k downloads1y agoHugging Face17mteb /MIRACLReranking MIRACLReranking An MTEB dataset Massive Text Embedding Benchmark MIRACL (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual retrieval dataset that focuses on search across 18 different languages. Task category t2t Domains Encyclopaedic, Written Reference https://project-miracl.github.io/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/MIRACLReranking.texttext-ranking1M<n<10M0 likes1.8k downloads1y agoHugging Face18mirav /artistic-imagery-altcaptionstext1K<n<10K1 likes1.6k downloads3y agoHugging Face19leeaandrob /mirror-eduagarcia__CrawlPT_dedup CrawlPT (deduplicated) CrawlPT is a generic Portuguese corpus extracted from various web pages. This version is deduplicated using MinHash algorithm and Locality Sensitive Hashing, following the approach of Lee et al. (2022). The raw version is also available here. Dataset Details Dataset is composed by three corpora: brWaC, C100-PT, OSCAR-2301. brWaC: a web corpus for Brazilian Portuguese from 120,000 different websites. C100-PT: Portuguese subset from CC-100.… See the full description on the dataset page: https://huggingface.co/datasets/leeaandrob/mirror-eduagarcia__CrawlPT_dedup.tabulartext-generation100M<n<1B0 likes1.4k downloads3mo agoHugging Face20mirror123 /ComPile Dataset Card for ComPile: A Large IR Dataset from Production Sources Changelog Release Programming Languages Description v1.0 C/C++, Rust, Swift, Julia Fine Tuning-scale dataset of 602GB of deduplicated LLVM (bitcode) IR Dataset Summary ComPile contains over 2.7TB of permissively-licensed source code compiled to (textual) LLVM intermediate representation (IR) covering C/C++, Rust, Swift, and Julia. The dataset was created by hooking into LLVM… See the full description on the dataset page: https://huggingface.co/datasets/mirror123/ComPile.texttext-generation100K<n<1M0 likes1.4k downloads8mo agoHugging Face21tomaarsen /miriad-4.4M-split MIRIAD 4.4M, split MIRIAD reformatted for training retrieval models: train, eval and test splits, and two subsets depending on what you want the model to retrieve. subset columns use it to retrieve default question, passage_text the source passage a question was generated from (averaging 941 tokens) question-answer question, answer the generated answer to a question (much shorter) split rows train 4,467,542 eval 10,000 test 10,000 [!TIP]… See the full description on the dataset page: https://huggingface.co/datasets/tomaarsen/miriad-4.4M-split.texttext-retrieval1M<n<10M5 likes1.3k downloads28d agoHugging Face22nvidia /miracl-vision MIRACL-VISION MIRACL-VISION is a multilingual visual retrieval dataset for 18 different languages. It is an extension of MIRACL, a popular text-only multilingual retrieval dataset. The dataset contains user questions, images of Wikipedia articles and annotations, which article can answer a user question. There are 7898 questions and 338734 images. More details can be found in the paper MIRACL-VISION: A Large, multilingual, visual document retrieval benchmark. This dataset is ready… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/miracl-vision.image100K<n<1M13 likes1.2k downloads1y agoHugging Face23MIRA-Lab /ChiPBench-D ChiPBench-D ChiPBench:Benchmarking End-to-End Performance of AI-based Chip Placement Algorithms Chip placement is a critical step in the Electronic Design Automation (EDA) workflow, which aims to arrange chip modules on the canvas to optimize the performance, power, and area (PPA) metrics of final designs.Recent advances show great potential of AI-based algorithms in chip placement.However, due to the lengthy EDA workflow, evaluations of these algorithms often focus on intermediate… See the full description on the dataset page: https://huggingface.co/datasets/MIRA-Lab/ChiPBench-D.textn<1K2 likes1.2k downloads1y agoHugging Face24Yunncheng /Mirage-Test 🌊 Mirage-Test Dataset Mirage-Test is a modern test-only dataset for benchmarking AI-generated image detection models. It contains real (0_real) and fake (1_fake) images across five distinct content domains, designed to evaluate generalization across diverse visual semantics. The fake images are generated using state-of-the-art generative models specifically optimized for perceptual realism and visual fidelity. 📌 This dataset is for evaluation only. No training split is… See the full description on the dataset page: https://huggingface.co/datasets/Yunncheng/Mirage-Test.imageimage-classification10K<n<100K3 likes1.1k downloads10mo agoHugging Face25miriad /miriad-5.8M Dataset Summary MIRIAD is a curated million scale Medical Instruction and RetrIeval Dataset. It contains 5.8 million medical question-answer pairs, distilled from peer-reviewed biomedical literature using LLMs. MIRIAD provides structured, high-quality QA pairs, enabling diverse downstream tasks like RAG, medical retrieval, hallucination detection, and instruction tuning. The dataset was introduced in our arXiv preprint. To load the dataset, run: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/miriad/miriad-5.8M.text1M<n<10M67 likes1.1k downloads1y agoHugging Face26mirianfsilva /multi_prompt_mmlutext10M<n<100M0 likes1k downloads2y agoHugging Face27magic-mirror /tsdm-lossless-music-3-othertext10K<n<100K0 likes947 downloads2y agoHugging Face28MIRAGE-Benchmark /MIRAGE MIRAGE Benchmark Project Page | Paper | GitHub MIRAGE is a benchmark for multimodal expert-level reasoning and decision-making in consultative interaction settings, specifically designed for the agriculture domain. It captures the complexity of expert consultations by combining natural user queries, expert-authored responses, and image-based context. The benchmark spans diverse crop health, pest diagnosis, and crop management scenarios, including more than 7,000 unique biological… See the full description on the dataset page: https://huggingface.co/datasets/MIRAGE-Benchmark/MIRAGE.imageimage-text-to-text10K<n<100K3 likes928 downloads8mo agoHugging Face29brhkim /education_data_portal_mirror_2026q3 Education Data Portal — Parquet Mirror (2026Q3 · Portal v0.26.1) A complete mirror of the Urban Institute Education Data Portal datasets version 0.26.1, collected August 6, 2026, and converted from CSV to Apache Parquet format for efficient analytical use. Please note that the maintainers of this Huggingface Dataset have no affiliation with the Urban Institute or the Education Data Portal team. Huge appreciation for all they do -- if you use this mirror, please make sure to… See the full description on the dataset page: https://huggingface.co/datasets/brhkim/education_data_portal_mirror_2026q3.tabular1B<n<10B0 likes906 downloads2mo agoHugging Face30aipracticecafe-mirror /curated-danbooru-2026-256px-flux2-vaetabular100K<n<1M0 likes820 downloads27d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.