datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
coco-2017-mirror
COCO 2017 mirror
This is a just mirror of the raw COCO dataset files, for convenience. You have to download it using something like:
pip install huggingface_hub
huggingface-cli download --local-dir coco-2017 pcuenq/coco-2017-mirror
And then unzip the files before use.
miradordatacore1librivox-mirror
LibriVox Mirror
Fast, structured, continuously updated LibriVox audio mirror.
Current snapshot
Metric
Value
Published books
21,724
Published sections
493,186
Audio hours
132,549.7
Audio languages
86
Quarantined books
610
Last updated (UTC)
2026-09-21T15:06:16.273192Z
Audio by language
Language
Hours
English
131,600.3
German
417.0
Spanish
160.9
French
103.8
Portuguese
37.4
Polish
34.1
Dutch
25.8… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/librivox-mirror.community-pipelines-mirror
Community Pipeline Examples
For more information about community pipelines, please have a look at this issue.
Community pipeline examples consist pipelines that have been added by the community.
Please have a look at the following tables to get an overview of all community examples. Click on the Code Example to get a copy-and-paste ready code example that you can try out.
If a community pipeline doesn't work as expected, please open an issue and ping the author on it.
Please… See the full description on the dataset page: https://huggingface.co/datasets/diffusers/community-pipelines-mirror.China-Building-Footprints-CMAB-Mirror
Origin Data
@misc{Zhang2025CMAB,
author = {Zhang, Yecheng and Zhao, Huimin and Long, Ying},
title = {{CMAB-The World's First National-Scale Multi-Attribute Building Dataset}},
year = {2025},
month = apr,
publisher = {figshare},
doi = {10.6084/m9.figshare.27992417},
url = {https://doi.org/10.6084/m9.figshare.27992417},
howpublished = {dataset}
}
Paper
@article{Zhang2025SciData,
author = {Zhang, Y. and… See the full description on the dataset page: https://huggingface.co/datasets/DannHiroaki/China-Building-Footprints-CMAB-Mirror.MIRACLRetrievalHardNegatives
MIRACLRetrievalHardNegatives
An MTEB dataset
Massive Text Embedding Benchmark
MIRACL (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual retrieval dataset that focuses on search across 18 different languages. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct.
Task category
t2t
Domains
Encyclopaedic, Written
Reference
http://miracl.ai/… See the full description on the dataset page: https://huggingface.co/datasets/mteb/MIRACLRetrievalHardNegatives.psg-audio-v3-unofficial-mirror
PSG-Audio v3 — Unofficial Complete Mirror
Unofficial complete mirror of the publicly released PSG-Audio Version 3 dataset.
This repository preserves the original files without modification and provides a reliable, high-speed mirror through the Hugging Face Hub for the research community.
Overview
PSG-Audio v3 is one of the largest publicly available multimodal sleep datasets, combining overnight clinical polysomnography (PSG) with synchronized environmental… See the full description on the dataset page: https://huggingface.co/datasets/Samuelsantos777/psg-audio-v3-unofficial-mirror.mirador-offloadESL-Bench
ESL-bench
ESL-bench (Event-driven Synthetic Longitudinal Benchmark) is a virtual health user dataset for evaluating AI health assistants. Each virtual user contains a complete health profile, event timeline, clinical exam data, and knowledge-graph-grounded evaluation queries, designed for use with the Mirobody-Eval framework.
⚠️ Research use only. Outputs are synthetic and intended for benchmarking AI agents. They should not be used for diagnosis or treatment decisions.… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/ESL-Bench.miracl
Dataset Card for MIRACL
This is a reformatting of the MIRACL dataset used to train the BGE-M3 model. See the full BGE-M3 dataset in Shitao/bge-m3-data.
Dataset Subsets
...-triplet subset
Columns: "anchor", "positive", "negative"
Column types: str, str, str
Examples:{
'anchor': '月球到地球的距离是多少?',
'positive': '月球距離\n月球距離 (LD) 是天文學上從地球到月球的距離,從地球到月球的平均距離是384,401公里 (238,856英里)。因為月球在橢圓軌道上運動,實際的距離隨時都在變化著。',
'negative':… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/miracl.LPNSR
LPNSR Dataset
This repository contains the evaluation datasets and testing data associated with the paper LPNSR: Optimal Noise-Guided Diffusion Image Super-Resolution Via Learnable Noise Prediction.
Project Links
Paper: arXiv:2603.21045
GitHub Repository: Faze-Hsw/LPNSR
Dataset Description
This dataset collection is used to evaluate image super-resolution models on both synthetic and complex real-world degradations. It contains pairs of Low-Quality (LQ) and… See the full description on the dataset page: https://huggingface.co/datasets/mirpri/LPNSR.model-mirror-17miracl-corpus
Dataset Card for MIRACL Corpus
MIRACL 🌍🙌🌏 (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual retrieval dataset that focuses on search across 18 different languages, which collectively encompass over three billion native speakers around the world.
This dataset contains the collection data of the 16 "known languages". The remaining 2 "surprise languages" will not be released until later.
The corpus for each language is prepared from a Wikipedia… See the full description on the dataset page: https://huggingface.co/datasets/miracl/miracl-corpus.MirrorPPR47MPaper: https://arxiv.org/abs/2606.29308
MedHall-Bench
MedHall-Bench
MedHall-Bench is a field-grounded hallucination detection benchmark for medical AI assistants. It decomposes each clinical response into verifiable structured fields (dose value, unit, reference range, ICD/LOINC code, entity relation, ...) and evaluates AI outputs via per-field programmatic matching in addition to sentence-level LLM-as-Judge. Designed for use with the HolyEval framework.
⚠️ Research use only. Content is for benchmarking AI agents and should not be… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/MedHall-Bench.MedHarm-Bench
MedHarm-Bench
MedHarm-Bench is a red-team compliance benchmark for health-management AI assistants. It uses natural-sounding patient questions that bait the assistant into crossing medical safety boundaries, then scores each response against compliance red lines. Designed for use with the HolyEval framework.
⚠️ Research use only. Questions are designed to elicit unsafe behavior for benchmarking purposes and should not be used for diagnosis or treatment decisions.… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/MedHarm-Bench.MiroFlow-BenchmarksThese are the benchmarking datasets used for MiroFlow Framework. More information: https://github.com/MiroMindAI/MiroThinker
deep1b-mirror
Deep1B — mirror
Mirror of the Deep1B billion-scale ANN benchmark (1,000,000,000 × 96-dim
float32 deep image descriptors) and its 10,000-query public test set,
originally published by Yandex Research and hosted at
https://storage.yandexcloud.net/yandex-research/ann-datasets/DEEP/.
This is a verbatim byte-for-byte copy intended for ANN-benchmark
reproducibility. Original authors / license: Yandex Research; please
cite their work (Babenko & Lempitsky, "Efficient Indexing of… See the full description on the dataset page: https://huggingface.co/datasets/suchun/deep1b-mirror.China-Building-Footprints-CMAB-Mirror
Origin Data
@misc{Zhang2025CMAB,
author = {Zhang, Yecheng and Zhao, Huimin and Long, Ying},
title = {{CMAB-The World's First National-Scale Multi-Attribute Building Dataset}},
year = {2025},
month = apr,
publisher = {figshare},
doi = {10.6084/m9.figshare.27992417},
url = {https://doi.org/10.6084/m9.figshare.27992417},
howpublished = {dataset}
}
Paper
@article{Zhang2025SciData,
author = {Zhang, Y. and… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/China-Building-Footprints-CMAB-Mirror.councilof-ai-mirror
Council of AI — public evidence mirror
Mirrored from https://councilof.ai at 2026-09-22T09:37:30Z.
councilof.ai is authoritative. This mirror exists for one measured reason: the origin answers HTTP 403 (Cloudflare error 1010) to a plain Python or Perl client, on the API as well as the site, so a machine consumer following our own published instructions is refused. Hugging Face serves those same clients.
Every file carries the URL it came from and the sha256 of the bytes as… See the full description on the dataset page: https://huggingface.co/datasets/csoai/councilof-ai-mirror.model-mirror-1MIRACLRetrieval
MIRACLRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
MIRACL (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual retrieval dataset that focuses on search across 18 different languages.
Task category
t2t
Domains
Encyclopaedic, Written
Reference
http://miracl.ai/
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/MIRACLRetrieval.model-mirror-15mirage18k
[IROS 2026] Mirage 18k: Dataset for Glass Segmentation & Depth Estimation
Mirage 18k is a novel, multi-task dataset comprising 18,353 manually annotated images across 38 unique indoor scenes, designed specifically for joint glass segmentation and glass-aware monocular depth estimation in robotics.
It contains diverse real-world glass structures (indoor panes, frosted doors, windows, clear doors) with severe background clutter, saliency, and dynamic obstacles.
Model Checkpoint:… See the full description on the dataset page: https://huggingface.co/datasets/rtarun1/mirage18k.miracl
Dataset Card for MIRACL (Topics and Qrels)
Dataset Description
Homepage |
Repository: |
Paper |
ArXiv
MIRACL 🌍🙌🌏 (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual retrieval dataset that focuses on search across 18 different languages, which collectively encompass over three billion native speakers around the world.
This dataset contains the collection data of the 16 "known languages". The remaining 2 "surprise languages" will not… See the full description on the dataset page: https://huggingface.co/datasets/miracl/miracl.model-mirror-29nerf-synthetic-mirrormodel-mirror-12cascade-testnet-mirroranime-syntheticsMostly unfiltered anime-style images generated by various text to image models, collected from various sources (some were submitted for inclusion by their creators).
Includes a subset of p1atdev/niji-v5, albeit captioned differently than the source.
Contains 2224 image & caption pairs.
As it is unfiltered, some adult content may be included.
Captions may not be completely accurate.
If you wish to submit content, do it as a pull request.
