datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
librivox-mirror
LibriVox Mirror
Fast, structured, continuously updated LibriVox audio mirror.
Current snapshot
Metric
Value
Published books
21,734
Published sections
493,396
Audio hours
132,613.0
Audio languages
86
Quarantined books
603
Last updated (UTC)
2026-09-24T13:43:06.550876Z
Audio by language
Language
Hours
English
131,663.6
German
417.0
Spanish
160.9
French
103.8
Portuguese
37.4
Polish
34.1
Dutch
25.8… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/librivox-mirror.jalandhary_asrJalandhary dataset is created using whisper model for STT and TTS. Some audios are ommited due to issues while trimming them. If there are some isues
in the dataset or audio not matching the text you can start a discussion or ping me to correcting it.
example_mmdata_mnbvc
mnbvc mm dataset v2.1
MNBVC 多模态语料数据格式。原链接:https://huggingface.co/datasets/wanng/example_mmdata_mnbvc
参考实现:mm_template_mnbvc
的 mmdata_block.BLOCK_SCHEMA。schema 以那份代码为准,这个数据集是它的示例产物。
字段
字段名称
类型
字段说明
可选
实体ID
string
数据的唯一标识符。用于在数据集中确定是哪一条数据。在单个数据集中确定一条数据的实体对象。
必选
md5
string
内容的 md5,用于去重与完整性校验
必选
块ID
int32
一个实体对象内的标识符。用于确定一条数据内的一个部分数据。parquet 行的最小单元。
必选
块类型
string
用于保存块的类别。类别的含义为「模态」。取值见下
必选
扩展字段
string
用于保存块的元信息。为可以被成功 load 的 json 字符串。后期可继续扩展
必选… See the full description on the dataset page: https://huggingface.co/datasets/miracleyin/example_mmdata_mnbvc.phoneme_asrThis dataset contains the phonetic transcriptions of audios as well as English transcripts. Phonetic transcriptions are based on the g2p model. It can be used to train phoneme recognition
model using wav2vec2.
minds14-mirrorMINDS-14 is training and evaluation resource for intent
detection task with spoken data. It covers 14
intents extracted from a commercial system
in the e-banking domain, associated with spoken examples in 14 diverse language varieties.
