datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
librivox-mirror
LibriVox Mirror
Fast, structured, continuously updated LibriVox audio mirror.
Current snapshot
Metric
Value
Published books
21,725
Published sections
493,206
Audio hours
132,555.1
Audio languages
86
Quarantined books
609
Last updated (UTC)
2026-09-22T13:24:09.765701Z
Audio by language
Language
Hours
English
131,605.7
German
417.0
Spanish
160.9
French
103.8
Portuguese
37.4
Polish
34.1
Dutch
25.8… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/librivox-mirror.China-Building-Footprints-CMAB-Mirror
Origin Data
@misc{Zhang2025CMAB,
author = {Zhang, Yecheng and Zhao, Huimin and Long, Ying},
title = {{CMAB-The World's First National-Scale Multi-Attribute Building Dataset}},
year = {2025},
month = apr,
publisher = {figshare},
doi = {10.6084/m9.figshare.27992417},
url = {https://doi.org/10.6084/m9.figshare.27992417},
howpublished = {dataset}
}
Paper
@article{Zhang2025SciData,
author = {Zhang, Y. and… See the full description on the dataset page: https://huggingface.co/datasets/DannHiroaki/China-Building-Footprints-CMAB-Mirror.MIRACLRetrievalHardNegatives
MIRACLRetrievalHardNegatives
An MTEB dataset
Massive Text Embedding Benchmark
MIRACL (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual retrieval dataset that focuses on search across 18 different languages. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct.
Task category
t2t
Domains
Encyclopaedic, Written
Reference
http://miracl.ai/… See the full description on the dataset page: https://huggingface.co/datasets/mteb/MIRACLRetrievalHardNegatives.psg-audio-v3-unofficial-mirror
PSG-Audio v3 — Unofficial Complete Mirror
Unofficial complete mirror of the publicly released PSG-Audio Version 3 dataset.
This repository preserves the original files without modification and provides a reliable, high-speed mirror through the Hugging Face Hub for the research community.
Overview
PSG-Audio v3 is one of the largest publicly available multimodal sleep datasets, combining overnight clinical polysomnography (PSG) with synchronized environmental… See the full description on the dataset page: https://huggingface.co/datasets/Samuelsantos777/psg-audio-v3-unofficial-mirror.mirador-offloadESL-Bench
ESL-bench
ESL-bench (Event-driven Synthetic Longitudinal Benchmark) is a virtual health user dataset for evaluating AI health assistants. Each virtual user contains a complete health profile, event timeline, clinical exam data, and knowledge-graph-grounded evaluation queries, designed for use with the Mirobody-Eval framework.
⚠️ Research use only. Outputs are synthetic and intended for benchmarking AI agents. They should not be used for diagnosis or treatment decisions.… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/ESL-Bench.miracl
Dataset Card for MIRACL
This is a reformatting of the MIRACL dataset used to train the BGE-M3 model. See the full BGE-M3 dataset in Shitao/bge-m3-data.
Dataset Subsets
...-triplet subset
Columns: "anchor", "positive", "negative"
Column types: str, str, str
Examples:{
'anchor': '月球到地球的距离是多少?',
'positive': '月球距離\n月球距離 (LD) 是天文學上從地球到月球的距離,從地球到月球的平均距離是384,401公里 (238,856英里)。因為月球在橢圓軌道上運動,實際的距離隨時都在變化著。',
'negative':… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/miracl.LPNSR
LPNSR Dataset
This repository contains the evaluation datasets and testing data associated with the paper LPNSR: Optimal Noise-Guided Diffusion Image Super-Resolution Via Learnable Noise Prediction.
Project Links
Paper: arXiv:2603.21045
GitHub Repository: Faze-Hsw/LPNSR
Dataset Description
This dataset collection is used to evaluate image super-resolution models on both synthetic and complex real-world degradations. It contains pairs of Low-Quality (LQ) and… See the full description on the dataset page: https://huggingface.co/datasets/mirpri/LPNSR.miracl-corpus
Dataset Card for MIRACL Corpus
MIRACL 🌍🙌🌏 (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual retrieval dataset that focuses on search across 18 different languages, which collectively encompass over three billion native speakers around the world.
This dataset contains the collection data of the 16 "known languages". The remaining 2 "surprise languages" will not be released until later.
The corpus for each language is prepared from a Wikipedia… See the full description on the dataset page: https://huggingface.co/datasets/miracl/miracl-corpus.MedHall-Bench
MedHall-Bench
MedHall-Bench is a field-grounded hallucination detection benchmark for medical AI assistants. It decomposes each clinical response into verifiable structured fields (dose value, unit, reference range, ICD/LOINC code, entity relation, ...) and evaluates AI outputs via per-field programmatic matching in addition to sentence-level LLM-as-Judge. Designed for use with the HolyEval framework.
⚠️ Research use only. Content is for benchmarking AI agents and should not be… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/MedHall-Bench.MedHarm-Bench
MedHarm-Bench
MedHarm-Bench is a red-team compliance benchmark for health-management AI assistants. It uses natural-sounding patient questions that bait the assistant into crossing medical safety boundaries, then scores each response against compliance red lines. Designed for use with the HolyEval framework.
⚠️ Research use only. Questions are designed to elicit unsafe behavior for benchmarking purposes and should not be used for diagnosis or treatment decisions.… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/MedHarm-Bench.China-Building-Footprints-CMAB-Mirror
Origin Data
@misc{Zhang2025CMAB,
author = {Zhang, Yecheng and Zhao, Huimin and Long, Ying},
title = {{CMAB-The World's First National-Scale Multi-Attribute Building Dataset}},
year = {2025},
month = apr,
publisher = {figshare},
doi = {10.6084/m9.figshare.27992417},
url = {https://doi.org/10.6084/m9.figshare.27992417},
howpublished = {dataset}
}
Paper
@article{Zhang2025SciData,
author = {Zhang, Y. and… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/China-Building-Footprints-CMAB-Mirror.MIRACLRetrieval
MIRACLRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
MIRACL (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual retrieval dataset that focuses on search across 18 different languages.
Task category
t2t
Domains
Encyclopaedic, Written
Reference
http://miracl.ai/
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/MIRACLRetrieval.mirage18k
[IROS 2026] Mirage 18k: Dataset for Glass Segmentation & Depth Estimation
Mirage 18k is a novel, multi-task dataset comprising 18,353 manually annotated images across 38 unique indoor scenes, designed specifically for joint glass segmentation and glass-aware monocular depth estimation in robotics.
It contains diverse real-world glass structures (indoor panes, frosted doors, windows, clear doors) with severe background clutter, saliency, and dynamic obstacles.
Model Checkpoint:… See the full description on the dataset page: https://huggingface.co/datasets/rtarun1/mirage18k.anime-syntheticsMostly unfiltered anime-style images generated by various text to image models, collected from various sources (some were submitted for inclusion by their creators).
Includes a subset of p1atdev/niji-v5, albeit captioned differently than the source.
Contains 2224 image & caption pairs.
As it is unfiltered, some adult content may be included.
Captions may not be completely accurate.
If you wish to submit content, do it as a pull request.
MIRACLRetrieval
MIRACLRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
MIRACL (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual retrieval dataset that focuses on search across 18 different languages.
Task category
t2t
Domains
Encyclopaedic, Written
Reference
http://miracl.ai/
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/RSamoed/MIRACLRetrieval.MIRACLReranking
MIRACLReranking
An MTEB dataset
Massive Text Embedding Benchmark
MIRACL (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual retrieval dataset that focuses on search across 18 different languages.
Task category
t2t
Domains
Encyclopaedic, Written
Reference
https://project-miracl.github.io/
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/MIRACLReranking.artistic-imagery-altcaptionsmirror-eduagarcia__CrawlPT_dedup
CrawlPT (deduplicated)
CrawlPT is a generic Portuguese corpus extracted from various web pages.
This version is deduplicated using MinHash algorithm and Locality Sensitive Hashing, following the approach of Lee et al. (2022).
The raw version is also available here.
Dataset Details
Dataset is composed by three corpora:
brWaC, C100-PT, OSCAR-2301.
brWaC: a web corpus for Brazilian Portuguese from 120,000 different websites.
C100-PT: Portuguese subset from CC-100.… See the full description on the dataset page: https://huggingface.co/datasets/leeaandrob/mirror-eduagarcia__CrawlPT_dedup.ComPile
Dataset Card for ComPile: A Large IR Dataset from Production Sources
Changelog
Release
Programming Languages
Description
v1.0
C/C++, Rust, Swift, Julia
Fine Tuning-scale dataset of 602GB of deduplicated LLVM (bitcode) IR
Dataset Summary
ComPile contains over 2.7TB of permissively-licensed source code compiled to (textual) LLVM
intermediate representation (IR) covering C/C++, Rust, Swift, and Julia.
The dataset was created by hooking into LLVM… See the full description on the dataset page: https://huggingface.co/datasets/mirror123/ComPile.miriad-4.4M-split
MIRIAD 4.4M, split
MIRIAD reformatted for training retrieval
models: train, eval and test splits, and two subsets depending on what you want the model to
retrieve.
subset
columns
use it to retrieve
default
question, passage_text
the source passage a question was generated from (averaging 941 tokens)
question-answer
question, answer
the generated answer to a question (much shorter)
split
rows
train
4,467,542
eval
10,000
test
10,000
[!TIP]… See the full description on the dataset page: https://huggingface.co/datasets/tomaarsen/miriad-4.4M-split.miracl-vision
MIRACL-VISION
MIRACL-VISION is a multilingual visual retrieval dataset for 18 different languages. It is an extension of MIRACL, a popular text-only multilingual retrieval dataset. The dataset contains user questions, images of Wikipedia articles and annotations, which article can answer a user question. There are 7898 questions and 338734 images. More details can be found in the paper MIRACL-VISION: A Large, multilingual, visual document retrieval benchmark.
This dataset is ready… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/miracl-vision.ChiPBench-D
ChiPBench-D
ChiPBench:Benchmarking End-to-End Performance of AI-based Chip Placement Algorithms
Chip placement is a critical step in the Electronic Design Automation (EDA) workflow, which aims to arrange chip modules on the canvas to optimize the performance, power, and area (PPA) metrics of final designs.Recent advances show great potential of AI-based algorithms in chip placement.However, due to the lengthy EDA workflow, evaluations of these algorithms often focus on intermediate… See the full description on the dataset page: https://huggingface.co/datasets/MIRA-Lab/ChiPBench-D.Mirage-Test
🌊 Mirage-Test Dataset
Mirage-Test is a modern test-only dataset for benchmarking AI-generated image detection models.
It contains real (0_real) and fake (1_fake) images across five distinct content domains, designed to evaluate generalization across diverse visual semantics.
The fake images are generated using state-of-the-art generative models specifically optimized for perceptual realism and visual fidelity.
📌 This dataset is for evaluation only. No training split is… See the full description on the dataset page: https://huggingface.co/datasets/Yunncheng/Mirage-Test.miriad-5.8M
Dataset Summary
MIRIAD is a curated million scale Medical Instruction and RetrIeval Dataset. It contains 5.8 million medical question-answer pairs, distilled from peer-reviewed biomedical literature using LLMs. MIRIAD provides structured, high-quality QA pairs, enabling diverse downstream tasks like RAG, medical retrieval, hallucination detection, and instruction tuning.
The dataset was introduced in our arXiv preprint.
To load the dataset, run:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/miriad/miriad-5.8M.multi_prompt_mmlutsdm-lossless-music-3-otherMIRAGE
MIRAGE Benchmark
Project Page | Paper | GitHub
MIRAGE is a benchmark for multimodal expert-level reasoning and decision-making in consultative interaction settings, specifically designed for the agriculture domain. It captures the complexity of expert consultations by combining natural user queries, expert-authored responses, and image-based context.
The benchmark spans diverse crop health, pest diagnosis, and crop management scenarios, including more than 7,000 unique biological… See the full description on the dataset page: https://huggingface.co/datasets/MIRAGE-Benchmark/MIRAGE.education_data_portal_mirror_2026q3
Education Data Portal — Parquet Mirror (2026Q3 · Portal v0.26.1)
A complete mirror of the Urban Institute Education Data Portal datasets version 0.26.1, collected August 6, 2026, and converted from CSV to Apache Parquet format for efficient analytical use. Please note that the maintainers of this Huggingface Dataset have no affiliation with the Urban Institute or the Education Data Portal team.
Huge appreciation for all they do -- if you use this mirror, please make sure to… See the full description on the dataset page: https://huggingface.co/datasets/brhkim/education_data_portal_mirror_2026q3.curated-danbooru-2026-256px-flux2-vae
