datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
librivox-mirror
LibriVox Mirror
Fast, structured, continuously updated LibriVox audio mirror.
Current snapshot
Metric
Value
Published books
21,725
Published sections
493,206
Audio hours
132,555.1
Audio languages
86
Quarantined books
609
Last updated (UTC)
2026-09-22T13:24:09.765701Z
Audio by language
Language
Hours
English
131,605.7
German
417.0
Spanish
160.9
French
103.8
Portuguese
37.4
Polish
34.1
Dutch
25.8… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/librivox-mirror.mirador-offloadmirror-eduagarcia__CrawlPT_dedup
CrawlPT (deduplicated)
CrawlPT is a generic Portuguese corpus extracted from various web pages.
This version is deduplicated using MinHash algorithm and Locality Sensitive Hashing, following the approach of Lee et al. (2022).
The raw version is also available here.
Dataset Details
Dataset is composed by three corpora:
brWaC, C100-PT, OSCAR-2301.
brWaC: a web corpus for Brazilian Portuguese from 120,000 different websites.
C100-PT: Portuguese subset from CC-100.… See the full description on the dataset page: https://huggingface.co/datasets/leeaandrob/mirror-eduagarcia__CrawlPT_dedup.miracl-vision
MIRACL-VISION
MIRACL-VISION is a multilingual visual retrieval dataset for 18 different languages. It is an extension of MIRACL, a popular text-only multilingual retrieval dataset. The dataset contains user questions, images of Wikipedia articles and annotations, which article can answer a user question. There are 7898 questions and 338734 images. More details can be found in the paper MIRACL-VISION: A Large, multilingual, visual document retrieval benchmark.
This dataset is ready… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/miracl-vision.education_data_portal_mirror_2026q3
Education Data Portal — Parquet Mirror (2026Q3 · Portal v0.26.1)
A complete mirror of the Urban Institute Education Data Portal datasets version 0.26.1, collected August 6, 2026, and converted from CSV to Apache Parquet format for efficient analytical use. Please note that the maintainers of this Huggingface Dataset have no affiliation with the Urban Institute or the Education Data Portal team.
Huge appreciation for all they do -- if you use this mirror, please make sure to… See the full description on the dataset page: https://huggingface.co/datasets/brhkim/education_data_portal_mirror_2026q3.curated-danbooru-2026-256px-flux2-vaeMiraData
MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions
Xuan Ju1*, Yiming Gao1*, Zhaoyang Zhang1*#, Ziyang Yuan1, Xintao Wang1, Ailing Zeng, Yu Xiong, Qiang Xu, Ying Shan1
1ARC Lab, Tencent PCG 2The Chinese University of Hong Kong *Equal Contribution #Project Lead
Introduction
Video datasets play a crucial role in video generation such as Sora.
However, existing text-video datasets often fall short when it comes to handling long video… See the full description on the dataset page: https://huggingface.co/datasets/TencentARC/MiraData.NSL-KDD
NSL-KDD
The data set is a data set that converts the arff File provided by the link into CSV and results.
The data set is personally stored by converting data to float64.
If you want to obtain additional original files, they are organized in the Original Directory in the repo.
Labels
The label of the data set is as follows.
#
Column
Non-Null
Count
Dtype
0
duration
151165
non-null
int64
1
protocol_type
151165
non-null
object
2
service
151165
non-null… See the full description on the dataset page: https://huggingface.co/datasets/Mireu-Lab/NSL-KDD.muchomusic
MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models
MuChoMusic is a benchmark designed to evaluate music understanding in multimodal language models focused on audio. It includes 1,187 multiple-choice questions validated by human annotators, based on 644 music tracks from two publicly available music datasets. These questions cover a wide variety of genres and assess knowledge and reasoning across several musical concepts and their cultural and functional… See the full description on the dataset page: https://huggingface.co/datasets/mulab-mir/muchomusic.vqav2-full-metadataUNSW-NB15
UNSW-NB15
This data is provided through the Train, Test CSV file provided by UNSW-NB15.
link
Labels
The label of the data set is as follows.
#
Column
Non-Null
Count
Dtype
0
id
82332
non-null
int64
1
dur
82332
non-null
float64
2
proto
82332
non-null
object
3
service
82332
non-null
object
4
state
82332
non-null
object
5
spkts
82332
non-null
int64
6
dpkts
82332
non-null
int64
7
sbytes
82332
non-null
int64
8
dbytes
82332
non-null
int64
9
rate… See the full description on the dataset page: https://huggingface.co/datasets/Mireu-Lab/UNSW-NB15.MiroVerse-v0.1mirror-eduagarcia__LegalPT_dedup
LegalPT (deduplicated)
LegalPT aggregates the maximum amount of publicly available legal data in Portuguese, drawing from varied sources including legislation, jurisprudence, legal articles, and government documents.
This version is deduplicated using MinHash algorithm and Locality Sensitive Hashing, following the approach of Lee et al. (2022).
The raw version is also available here.
Dataset Details
Dataset is composed by six corpora:
Ulysses-Tesemõ… See the full description on the dataset page: https://huggingface.co/datasets/leeaandrob/mirror-eduagarcia__LegalPT_dedup.mirth_lerobot
MIRTH Dataset
Multi-camera real-world manipulation demonstrations for history-aware Vision-Language-Action agents
The MIRTH dataset is a real-world robot manipulation dataset collected on a physical LeRobot platform. It contains synchronized main-camera and wrist-camera observations, robot proprioception, language instructions, and expert action trajectories for training and evaluating Vision-Language-Action (VLA) agents.
This release provides the same… See the full description on the dataset page: https://huggingface.co/datasets/Kiva12138/mirth_lerobot.MIRACLE
MIRACLE
MIRACLE is a multimodal benchmark dataset with image-based questions and model evaluation results.
This repository contains the test subset of the MIRACLE benchmark. The benchmark config provides the benchmark test split, and the model_results config provides test-set evaluation outputs for the included models.
Dataset Structure
The repository is organized as follows:
data/
test.parquet # Hugging Face loadable benchmark test split
test.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/queyuecanyang/MIRACLE.ovos-wake-word-bench-picovoice-smart-mirror
OVOS wake_word bench — picovoice-smart-mirror
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
Picovoice/wake-word-benchmark.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-smart-mirror.MIRAGE-CanaryDocs
MIRAGE CanaryDocs
MIRAGE CanaryDocs is an English synthetic enterprise-document dataset for structured privacy-unit,
canary, and ordered multi-chunk evaluation. It is the companion dataset for the EMNLP 2026 paper
When Metadata Remembers: Ordered Provenance Enables Document-Level Embedding Inversion.
Project documentation and schemas are also available in the
MIRAGE GitHub repository.
Dataset summary
The dataset contains complete synthetic documents, ordered token… See the full description on the dataset page: https://huggingface.co/datasets/LevenKoko/MIRAGE-CanaryDocs.agilex_push_mirror_surfaceThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "arx5_bimanual",
"total_episodes": 20,
"total_frames": 6503,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 25,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/agilex_push_mirror_surface.level12_rac_2_2026-02-07_and_mirThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 10770,
"total_frames": 26748966,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:10770"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/level12_rac_2_2026-02-07_and_mir.backln-guest-post-quality-public-mirror
Backln Guest Post Quality Public Mirror
Public-safe mirror for validating Hugging Face Dataset Viewer indexing and release gates. This dataset is not the private training corpus.
Full text, titles, and snippets are removed by default. The mirror keeps labels, coarse metadata, feature buckets, and hash prefixes so the public Hub can verify schema and distribution without exposing customer content.
Schema
label: one of published, manual_review, rejected.
source: coarse… See the full description on the dataset page: https://huggingface.co/datasets/driodnexus/backln-guest-post-quality-public-mirror.mir2023agilex_cover_mirror_jshThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "arx5_bimanual",
"total_episodes": 20,
"total_frames": 7153,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 25,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/agilex_cover_mirror_jsh.education_data_portal_mirror
⚠️ Frozen Vintage — Education Data Portal Parquet Mirror (v0.24.0, February 2026)
This repository is frozen and will receive no further updates. It is preserved
permanently as a snapshot for reproducibility. For current data, use the
successor mirror:
brhkim/education_data_portal_mirror_2026q3
(Education Data Portal v0.26.1, collected August 2026).
If you are reproducing an analysis that originally used this mirror, you are in the
right place — keep your scripts pointed here.… See the full description on the dataset page: https://huggingface.co/datasets/brhkim/education_data_portal_mirror.MirrorAPI-Training
MirrorAPI training dataset
This dataset contains the training data for MirrorAPI and MirrorAPI-Cache:
train_sft.json, train_cot.json, train_augment.json: The training data for MirrorAPI .
train_cache.json: The training data for MirrorAPI-Cache.
proofwriter-mirror
ProofWriter (The Mirror)
A typed, content-faithful mirror of AI2's ProofWriter dataset (release V2020.12.3),
derived from proofwriter-source. The JSON encoding is cleaned up: the id-keyed dicts
(triple1, Q3, …) become lists of structs that keep their id, every atom representation
is parsed into a typed {subject, relation, object, polarity} triple, and the closed enums
(answer, strategy) are typed. The content stays faithful — nothing renamed, no rows
dropped — and the recursive… See the full description on the dataset page: https://huggingface.co/datasets/arqa39/proofwriter-mirror.proofwriter-mirror
ProofWriter (The Mirror)
A typed, content-faithful mirror of AI2's ProofWriter dataset (release V2020.12.3),
derived from proofwriter-source. The JSON encoding is cleaned up: the id-keyed dicts
(triple1, Q3, …) become lists of structs that keep their id, every atom representation
is parsed into a typed {subject, relation, object, polarity} triple, and the closed enums
(answer, strategy) are typed. The content stays faithful — nothing renamed, no rows
dropped — and the recursive… See the full description on the dataset page: https://huggingface.co/datasets/rlhf-and-friends/proofwriter-mirror.03-09-Left-ReGrasp_mirrored_right_armThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aibot2",
"total_episodes": 57,
"total_frames": 4552,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:57"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/alphabot2/03-09-Left-ReGrasp_mirrored_right_arm.V4_pnp_combined_mirroredThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Screener2/V4_pnp_combined_mirrored.stackoverflowVQA
Dataset Card for "stackoverflowVQA"
More Information needed
mawi-https-flows-2025
