datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MLAAD
Introduction
Welcome to MLAAD: The Multi-Language Audio Anti-Spoofing Dataset -- a dataset to train, test and evaluate audio deepfake detection. See
the paper for more information.
License
MLAAD is published strictly for non-commercial academic research use, under the CC-BY-NC 4.0 license. Commercial use is not permitted.
Bibtex
If you use this dataset, please consider citing it as follows.
@article{muller2024mlaad,
title={MLAAD: The… See the full description on the dataset page: https://huggingface.co/datasets/mueller91/MLAAD.Daimon-Infinity
Daimon-Infinity mirror
This repository is a file-preserving mirror of
daimonrobotics/Daimon-Infinity on ModelScope.
Source and license
Upstream: daimonrobotics/Daimon-Infinity
License: CC BY-NC-SA 4.0
Attribution: Daimon Robotics / Daimon-Infinity
This mirror keeps the upstream directory layout and is distributed under the
same CC BY-NC-SA 4.0 license. No data is altered; files are transferred with
integrity checks supplied by ModelScope and the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/ml-resources/Daimon-Infinity.peoples_speech
Dataset Card for People's Speech
Dataset Summary
The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/peoples_speech.unsupervised_peoples_speech
Dataset Card for Unsupervised Peoples Speech
Dataset Description
Dataset Summary
The Unsupervised Peoples Speech Dataset is a compilation of audiofiles extracted from Archive.org that is licensed for academic and commercial usage under CC-BY and CC-BY-SA licenses. It includes more than one million hours of audio with a diverse set of speakers.
Point of Contact: MLCommons Datasets Discord
Dataset Structure
This dataset is a collection of audio… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/unsupervised_peoples_speech.speech-wikimedia
Dataset Card for Speech Wikimedia
Dataset Summary
The Speech Wikimedia Dataset is a compilation of audiofiles with transcriptions extracted from wikimedia commons that is licensed for academic and commercial usage under CC and Public domain. It includes 2,000+ hours of transcribed speech in different languages with a diverse set of speakers.
Each audiofile should have one or more transcriptions in different languages.
Transcription languages
English
German… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/speech-wikimedia.whisper_transcriptions.mls.wer_10.0mls_sidon
MLS-Sidon
Overview
This dataset is a cleansed version of Multilingual LibriSpeech (MLS) with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling.
The dataset is provided in WebDataset format for efficient large-scale training.
Source: Multilingual LibriSpeech
Languages: English, German, French, Spanish, Italian, Polish, Dutch, Portuguese
Format: WebDataset (.tar shards)
License: CC-BY-4.0
Dataset Structure
Each sample in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/mls_sidon.mls_eng
Dataset Card for English MLS
Dataset Summary
This is a streamable version of the English version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng.en_asr.mlsMLAAD-tiny
Welcome to MLAAD-tiny
MLAAD-tiny is a very small subset of the full MLAAD dataset, designed for education, prototyping, and debugging.
Many teaching environments (e.g. Colab, Kaggle, university notebooks -- se this notebook for example) impose strict storage limits, which makes large-scale audio deepfake datasets impractical to use. To address this, we provide MLAAD-tiny, a compact yet representative version of MLAAD.
Download
git lfs install
git clone… See the full description on the dataset page: https://huggingface.co/datasets/mueller91/MLAAD-tiny.mls_hq_urgent_track1mls_eng_10k
Dataset Summary
This is a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng_10k.whisper_transcriptions.mlsLHCP-ASR
LHCP-ASR
This dataset is another version of the LHCP-ASR corpus, an English speech dataset for narrow-domain ASR benchmarking in high-energy physics. Unlike the original distribution, which includes video, slides and text data, this version focuses entirely on audio-text pairs
DESCRIPTION
The speech data are 30 hours of LHCP plenary conference talks (2020, 2022) with manual (human) verbatim transcriptions and 205 hours of LHCP conference talks (2020-2022) with automatic… See the full description on the dataset page: https://huggingface.co/datasets/mllp/LHCP-ASR.MLS-Sidon-HQ
MLS-Sidon-HQ
The top-DNSMOS slice of sarulab-speech/mls_sidon —
Multilingual LibriSpeech (LibriVox) restored to 48 kHz by sarulab-speech/sidon-v0.1,
then scored clip-by-clip with DNSMOS P.835 and cut down to only the cleanest utterances.
All 1,473,249 source clips (6,143 h, 0.91 TB of FLAC) were scored;
170,226 (11.6%, 702 h) passed and are published here.
Filter
DNSMOS P.835 (speechmos, ONNX) on a single centred 10 s window at 16 kHz,
peak-normalised. A clip is… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/MLS-Sidon-HQ.DFADD_MLAAD_DiffSSD_VoxCeleb2cv_mls_psfb_fs3_92Data used to reproduce all of the experiments in the paper ; https://ieeexplore.ieee.org/document/10720758/
mls_dutchmls_it_pseudo_labelled-large-v3free-music-archive-retrieval
FMAR: A Dataset for Robust Song Identification
Authors: Ryan Lee, Yi-Chieh Chiu, Abhir Karande, Ayush Goyal, Harrison Pearl, Matthew Hong, Spencer Cobb
Overview
To improve copyright infringement detection, we introduce Free-Music-Archive-Retrieval (FMAR), a structured dataset designed to test a model's capability to identify songs based on 5-second clips, or queries. We create adversarial queries to replicate common strategies to evade copyright infringement detectors… See the full description on the dataset page: https://huggingface.co/datasets/ml-ryanlee/free-music-archive-retrieval.hubert_layer9-librispeech-asr100h
Dataset Card for "hubert_layer9-librispeech-asr100h"
More Information needed
mm_mls_englishwavlm-large_layer21-librispeech-asr100h
Dataset Card for "wavlm-large_layer21-librispeech-asr100h"
More Information needed
ML2021_ASR_ST
Dataset Card for "ML2021_ASR_ST"
This dataset contains the audio recordings, the transcriptions, and the English translation of the transcriptions of the Machine Learning Course in 2021 at National Taiwan Univeristy.
This can be used for domain-specific and code-switching ASR/Speech-to-text translation.
If you find this dataset useful, please consider to cite the following paper:
@inproceedings{yang2024investigating,
title={Investigating zero-shot generalizability on… See the full description on the dataset page: https://huggingface.co/datasets/ky552/ML2021_ASR_ST.voxpopuli-mls-de-descriptions
Natural Language Voice Descriptions of the VoxPopuli and MLS German Datasets
German read and parliamentary speech paired with its transcript, acoustic
measurements, discrete German descriptor tags, and a free-text German
description of the speaker's voice and recording conditions. The dataset is intended
for training description-conditioned TTS models such as
Parler-TTS.
The data was built as part of research work. It is a random subset of the
pooled German portions of VoxPopuli… See the full description on the dataset page: https://huggingface.co/datasets/leonhard-behr/voxpopuli-mls-de-descriptions.indic_tts_ml
Indic TTS Malayalam Speech Corpus
The Malayalam subset of Indic TTS Corpus, taken from
this Kaggle database. The corpus contains
one male and one female speaker, with a 2:1 ratio of samples due to missing files for the female speaker. The license is given
in the repository.
cv_mls_psfb_fs0_24Data used to reproduce all of the experiments in the paper ; https://ieeexplore.ieee.org/document/10720758/
metric-mamba-ml2021-hungyi-corpus
Dataset Card for "metric-mamba-ml2021-hungyi-corpus"
More Information needed
mls-hq-urgent-track1
Multilingual LibriSpeech HQ (MLS-HQ)
This is a mirror of the Multilingual LibriSpeech HQ (MLS-HQ) data used in URGENT 2025 Track 1.
The original files were converted from FLAC to Opus to reduce the size and accelerate streaming.
Sampling rate: 48 kHz (resampled from 44.1 kHz to support Opus format)
Channels: 1
Format: Opus
Splits:
spanish: 150 hours, 36031 utterances
german: 150 hours, 35890 utterances
french: 150 hours, 36078 utterances
License: CC0 1.0
Source:… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/mls-hq-urgent-track1.cv_mls_psfb_zero_syntheticData used to reproduce all of the experiments in the paper ; https://ieeexplore.ieee.org/document/10720758/
