datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
librivox-mirror
LibriVox Mirror
Fast, structured, continuously updated LibriVox audio mirror.
Current snapshot
Metric
Value
Published books
21,728
Published sections
493,262
Audio hours
132,580.7
Audio languages
86
Quarantined books
606
Last updated (UTC)
2026-09-23T13:55:55.728817Z
Audio by language
Language
Hours
English
131,631.3
German
417.0
Spanish
160.9
French
103.8
Portuguese
37.4
Polish
34.1
Dutch
25.8… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/librivox-mirror.psg-audio-v3-unofficial-mirror
PSG-Audio v3 — Unofficial Complete Mirror
Unofficial complete mirror of the publicly released PSG-Audio Version 3 dataset.
This repository preserves the original files without modification and provides a reliable, high-speed mirror through the Hugging Face Hub for the research community.
Overview
PSG-Audio v3 is one of the largest publicly available multimodal sleep datasets, combining overnight clinical polysomnography (PSG) with synchronized environmental… See the full description on the dataset page: https://huggingface.co/datasets/Samuelsantos777/psg-audio-v3-unofficial-mirror.ami-corpus-mirror
AMI Corpus Mirror
Mirror of the subset of the AMI Meeting Corpus used by
FluidAudio diarization benchmarks. Hosted here so CI and
local benchmark runs do not depend on the availability of the upstream groups.inf.ed.ac.uk server
(see FluidAudio#752).
Contents
annotations/ami_public_manual_1.6.2.zip — AMI public manual annotations v1.6.2
(repackaged from the official archive; identical content, including segments/, words/,
corpusResources/meetings.xml)… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/ami-corpus-mirror.Mir-1k-use-DJCM-trainingvocalset-mirrorMirror of https://zenodo.org/records/1442513.
See the original dataset creators,
Wilkins, J., Prem Seetharaman, Alison Wahl, & Bryan Pardo. (2018). VocalSet: A Singing Voice Dataset (1.2) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.1442513
schale-archive-mirror
Schale Archive
Overview
This repository contains Blue Archive game assets.
Known Missing Characters Spine
NPC
GSC President
Yume
Nao
Nagusa
Nyanten-maru
Pei
Reizyo
A.R.O.N.A
How to Properly Import Character to Spine Editor
First, you need Spine Editor at least v3.8.x and later.
Create a new project.
Delete the default skeleton in Hierarchy on the right.
Click Spine logo on top left, then select Import Data.
Select… See the full description on the dataset page: https://huggingface.co/datasets/Hutao0514/schale-archive-mirror.musan-mirrorMiracle-ConversationThis is the dataset curated from ChatGPT with personalized prompt from Our EMNLP2023-findings Miracle
We offer three personality aspects:
'a' = attitude (positive/negative)
'l' = language style (lyrical/plain)
'm' = mental characteristics (critical/emotional)
audio-music-mir-post-public
audio-music-mir-post-public
Music information retrieval and tagging annotations: genre (FMA), instrument family (NSynth × 3, Medley-solos-DB), social tags (MagnaTagATune via LLARK), and large-scale Music4All metadata. Foundation for music understanding heads in audio LLMs.
Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py after fetching to rewrite the JSONL audio_path fields with absolute local… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-music-mir-post-public.jalandhary_asrJalandhary dataset is created using whisper model for STT and TTS. Some audios are ommited due to issues while trimming them. If there are some isues
in the dataset or audio not matching the text you can start a discussion or ping me to correcting it.
example_mmdata_mnbvc
mnbvc mm dataset v2.1
MNBVC 多模态语料数据格式。原链接:https://huggingface.co/datasets/wanng/example_mmdata_mnbvc
参考实现:mm_template_mnbvc
的 mmdata_block.BLOCK_SCHEMA。schema 以那份代码为准,这个数据集是它的示例产物。
字段
字段名称
类型
字段说明
可选
实体ID
string
数据的唯一标识符。用于在数据集中确定是哪一条数据。在单个数据集中确定一条数据的实体对象。
必选
md5
string
内容的 md5,用于去重与完整性校验
必选
块ID
int32
一个实体对象内的标识符。用于确定一条数据内的一个部分数据。parquet 行的最小单元。
必选
块类型
string
用于保存块的类别。类别的含义为「模态」。取值见下
必选
扩展字段
string
用于保存块的元信息。为可以被成功 load 的 json 字符串。后期可继续扩展
必选… See the full description on the dataset page: https://huggingface.co/datasets/miracleyin/example_mmdata_mnbvc.kids_phoneme_asr
Dataset Card for "kids_phoneme_asr"
More Information needed
kids_phoneme_md
Dataset Card for "kids_phoneme_md"
More Information needed
minoricloud-galgame-vocal-mp3-collection-1989-2023StressDetection_MIRSD
Dataset Card for "stress_dection_MIR_SD"
More Information needed
EmoV_DB-mirror3rd party mirror of https://github.com/numediart/EmoV-DB
Please see original license at https://github.com/numediart/EmoV-DB/blob/master/LICENSE.md
@article{adigwe2018emotional,
title={The emotional voices database: Towards controlling the emotion dimension in voice generation systems},
author={Adigwe, Adaeze and Tits, No{'e} and Haddad, Kevin El and Ostadabbas, Sarah and Dutoit, Thierry},
journal={arXiv preprint arXiv:1806.09514},
year={2018}
}
ristey_mirwari
Riste Mirwari — audiobook speech (ristey_mirwari)
31,535 transcribed speech segments cut from a reading of the book ڕشتەی مرواری
(Rishtey Mirwari), with a speaker and gender column alongside each segment.
At a glance
Rows
31,535 (single train split)
Columns
audio, file, transcription, speaker, gender, book, is_gold_transcript
Parquet on disk
4.99 GB across 11 shards (5.50 GB uncompressed)
Row groups
29 per shard (1,000 rows per group)
Audio… See the full description on the dataset page: https://huggingface.co/datasets/razhan/ristey_mirwari.mira-invocation-audio
Mira invocation audio
An exact snapshot of the active audio dataset used by the Mira Whisper and HuBERT
invocation classifiers. It includes 167 derivatives from one real microphone training session.
The label is 1 when the glasses wearer directly addresses the Mira assistant,
and 0 for incidental mentions, other conversation, or background audio.
Only the wearer's direct address should activate Mira. Synthetic positives assume
the synthetic speaker is the wearer; the classifiers… See the full description on the dataset page: https://huggingface.co/datasets/DhawalM/mira-invocation-audio.SAVEE-mirror3rd party mirror of https://www.kaggle.com/datasets/ejlok1/surrey-audiovisual-expressed-emotion-savee/data
Marked as "Data files © Original Authors" and shared on Kaggle datasets
ns-urdu-datasetphoneme_asrThis dataset contains the phonetic transcriptions of audios as well as English transcripts. Phonetic transcriptions are based on the g2p model. It can be used to train phoneme recognition
model using wav2vec2.
StressDetection_MIRSD_TTSTESS-mirrorAttribution-NonCommercial-NoDerivatives 4.0 by the original authors, shared on Kaggle first here: https://www.kaggle.com/datasets/ejlok1/toronto-emotional-speech-set-tess
blossoming-white-mirror
Blossoming White Mirror
License: Artistic-2.0Dataset ID: smolboy/blossoming-white-mirror
Overview
This dataset contains a complete science-fiction novel and associated media assets, structured in four narrative phases:
Political Thriller
Techno-Mystery
Mythic Science Fiction
Spiritual Metaphysics
It is provided for reading, audio playback, and research into narrative structure and multimedia processing.
Files
File
Type
Notes
Novel.pdf
PDF… See the full description on the dataset page: https://huggingface.co/datasets/smolboy/blossoming-white-mirror.tech-miracurolcasia-mirrorStressDetection_MIRSD_TTSsmart-turn-hindi-dataset-14kjamendo-qa-mirror
Jamendo-QA: A Large-Scale Music Question Answering Dataset
Jamendo-QA is a large-scale dataset designed for music-related question answering (Music-QA) research.It is built upon the Jamendo Music collection and supports research in music knowledge QA, audio-text multimodal learning, and music information retrieval (MIR).
📊 Dataset Summary
Item
Description
Name
Jamendo-QA
Domain
Music, Question Answering, MIR
Languages
English
Tasks
Question… See the full description on the dataset page: https://huggingface.co/datasets/Archit00/jamendo-qa-mirror.nptel_en_with_gender_and_speaker_classification
