mir
Datasets
All datasets matching “mir”coco-2017-mirror
COCO 2017 mirror
This is a just mirror of the raw COCO dataset files, for convenience. You have to download it using something like:
pip install huggingface_hub
huggingface-cli download --local-dir coco-2017 pcuenq/coco-2017-mirror
And then unzip the files before use.
miradordatacore1librivox-mirror
LibriVox Mirror
Fast, structured, continuously updated LibriVox audio mirror.
Current snapshot
Metric
Value
Published books
21,724
Published sections
493,186
Audio hours
132,549.7
Audio languages
86
Quarantined books
610
Last updated (UTC)
2026-09-21T15:06:16.273192Z
Audio by language
Language
Hours
English
131,600.3
German
417.0
Spanish
160.9
French
103.8
Portuguese
37.4
Polish
34.1
Dutch
25.8… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/librivox-mirror.community-pipelines-mirror
Community Pipeline Examples
For more information about community pipelines, please have a look at this issue.
Community pipeline examples consist pipelines that have been added by the community.
Please have a look at the following tables to get an overview of all community examples. Click on the Code Example to get a copy-and-paste ready code example that you can try out.
If a community pipeline doesn't work as expected, please open an issue and ping the author on it.
Please… See the full description on the dataset page: https://huggingface.co/datasets/diffusers/community-pipelines-mirror.China-Building-Footprints-CMAB-Mirror
Origin Data
@misc{Zhang2025CMAB,
author = {Zhang, Yecheng and Zhao, Huimin and Long, Ying},
title = {{CMAB-The World's First National-Scale Multi-Attribute Building Dataset}},
year = {2025},
month = apr,
publisher = {figshare},
doi = {10.6084/m9.figshare.27992417},
url = {https://doi.org/10.6084/m9.figshare.27992417},
howpublished = {dataset}
}
Paper
@article{Zhang2025SciData,
author = {Zhang, Y. and… See the full description on the dataset page: https://huggingface.co/datasets/DannHiroaki/China-Building-Footprints-CMAB-Mirror.MIRACLRetrievalHardNegatives
MIRACLRetrievalHardNegatives
An MTEB dataset
Massive Text Embedding Benchmark
MIRACL (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual retrieval dataset that focuses on search across 18 different languages. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct.
Task category
t2t
Domains
Encyclopaedic, Written
Reference
http://miracl.ai/… See the full description on the dataset page: https://huggingface.co/datasets/mteb/MIRACLRetrievalHardNegatives.

