CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mueller91 /MLAADgated Introduction Welcome to MLAAD: The Multi-Language Audio Anti-Spoofing Dataset -- a dataset to train, test and evaluate audio deepfake detection. See the paper for more information. License MLAAD is published strictly for non-commercial academic research use, under the CC-BY-NC 4.0 license. Commercial use is not permitted. Bibtex If you use this dataset, please consider citing it as follows. @article{muller2024mlaad, title={MLAAD: The… See the full description on the dataset page: https://huggingface.co/datasets/mueller91/MLAAD.audioaudio-classification100K<n<1M46 likes58k downloads26d agoHugging Face02ml-resources /Daimon-Infinity Daimon-Infinity mirror This repository is a file-preserving mirror of daimonrobotics/Daimon-Infinity on ModelScope. Source and license Upstream: daimonrobotics/Daimon-Infinity License: CC BY-NC-SA 4.0 Attribution: Daimon Robotics / Daimon-Infinity This mirror keeps the upstream directory layout and is distributed under the same CC BY-NC-SA 4.0 license. No data is altered; files are transferred with integrity checks supplied by ModelScope and the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/ml-resources/Daimon-Infinity.audio1K<n<10K2 likes39k downloads19d agoHugging Face03MLCommons /peoples_speech Dataset Card for People's Speech Dataset Summary The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/peoples_speech.audioautomatic-speech-recognition1M<n<10M285 likes39k downloads2y agoHugging Face04MLCommons /unsupervised_peoples_speech Dataset Card for Unsupervised Peoples Speech Dataset Description Dataset Summary The Unsupervised Peoples Speech Dataset is a compilation of audiofiles extracted from Archive.org that is licensed for academic and commercial usage under CC-BY and CC-BY-SA licenses. It includes more than one million hours of audio with a diverse set of speakers. Point of Contact: MLCommons Datasets Discord Dataset Structure This dataset is a collection of audio… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/unsupervised_peoples_speech.audioautomatic-speech-recognition81 likes26k downloads2y agoHugging Face05MLCommons /speech-wikimedia Dataset Card for Speech Wikimedia Dataset Summary The Speech Wikimedia Dataset is a compilation of audiofiles with transcriptions extracted from wikimedia commons that is licensed for academic and commercial usage under CC and Public domain. It includes 2,000+ hours of transcribed speech in different languages with a diverse set of speakers. Each audiofile should have one or more transcriptions in different languages. Transcription languages English German… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/speech-wikimedia.audion<1K14 likes23k downloads3y agoHugging Face06japanese-asr /whisper_transcriptions.mls.wer_10.0audio1M<n<10M2 likes16k downloads2y agoHugging Face07sarulab-speech /mls_sidon MLS-Sidon Overview This dataset is a cleansed version of Multilingual LibriSpeech (MLS) with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling. The dataset is provided in WebDataset format for efficient large-scale training. Source: Multilingual LibriSpeech Languages: English, German, French, Spanish, Italian, Polish, Dutch, Portuguese Format: WebDataset (.tar shards) License: CC-BY-4.0 Dataset Structure Each sample in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/mls_sidon.audiotext-to-speech10M<n<100M11 likes8.9k downloads1y agoHugging Face08parler-tts /mls_eng Dataset Card for English MLS Dataset Summary This is a streamable version of the English version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng.audioautomatic-speech-recognition10M<n<100M40 likes3.5k downloads2y agoHugging Face09japanese-asr /en_asr.mlsaudio10M<n<100M3 likes3.3k downloads2y agoHugging Face10mueller91 /MLAAD-tiny Welcome to MLAAD-tiny MLAAD-tiny is a very small subset of the full MLAAD dataset, designed for education, prototyping, and debugging. Many teaching environments (e.g. Colab, Kaggle, university notebooks -- se this notebook for example) impose strict storage limits, which makes large-scale audio deepfake datasets impractical to use. To address this, we provide MLAAD-tiny, a compact yet representative version of MLAAD. Download git lfs install git clone… See the full description on the dataset page: https://huggingface.co/datasets/mueller91/MLAAD-tiny.audioaudio-classification10K<n<100K3 likes3.1k downloads4mo agoHugging Face11kohei0209 /mls_hq_urgent_track1audio100K<n<1M0 likes1.7k downloads2y agoHugging Face12parler-tts /mls_eng_10k Dataset Summary This is a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng_10k.audioautomatic-speech-recognition1M<n<10M31 likes1.1k downloads2y agoHugging Face13japanese-asr /whisper_transcriptions.mlsaudio10M<n<100M1 likes966 downloads2y agoHugging Face14mllp /LHCP-ASR LHCP-ASR This dataset is another version of the LHCP-ASR corpus, an English speech dataset for narrow-domain ASR benchmarking in high-energy physics. Unlike the original distribution, which includes video, slides and text data, this version focuses entirely on audio-text pairs DESCRIPTION The speech data are 30 hours of LHCP plenary conference talks (2020, 2022) with manual (human) verbatim transcriptions and 205 hours of LHCP conference talks (2020-2022) with automatic… See the full description on the dataset page: https://huggingface.co/datasets/mllp/LHCP-ASR.audioautomatic-speech-recognition100K<n<1M0 likes772 downloads5mo agoHugging Face15Scicom-intl /MLS-Sidon-HQ MLS-Sidon-HQ The top-DNSMOS slice of sarulab-speech/mls_sidon — Multilingual LibriSpeech (LibriVox) restored to 48 kHz by sarulab-speech/sidon-v0.1, then scored clip-by-clip with DNSMOS P.835 and cut down to only the cleanest utterances. All 1,473,249 source clips (6,143 h, 0.91 TB of FLAC) were scored; 170,226 (11.6%, 702 h) passed and are published here. Filter DNSMOS P.835 (speechmos, ONNX) on a single centred 10 s window at 16 kHz, peak-normalised. A clip is… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/MLS-Sidon-HQ.audio100K<n<1M0 likes643 downloads2mo agoHugging Face16cmu-mlsp /DFADD_MLAAD_DiffSSD_VoxCeleb2audio1M<n<10M1 likes631 downloads1y agoHugging Face17remynd /cv_mls_psfb_fs3_92Data used to reproduce all of the experiments in the paper ; https://ieeexplore.ieee.org/document/10720758/ audio100K<n<1M0 likes595 downloads1y agoHugging Face18AdoCleanCode /mls_dutchaudio100K<n<1M1 likes588 downloads7mo agoHugging Face19deepdml /mls_it_pseudo_labelled-large-v3audio10K<n<100K0 likes576 downloads2y agoHugging Face20ml-ryanlee /free-music-archive-retrieval FMAR: A Dataset for Robust Song Identification Authors: Ryan Lee, Yi-Chieh Chiu, Abhir Karande, Ayush Goyal, Harrison Pearl, Matthew Hong, Spencer Cobb Overview To improve copyright infringement detection, we introduce Free-Music-Archive-Retrieval (FMAR), a structured dataset designed to test a model's capability to identify songs based on 5-second clips, or queries. We create adversarial queries to replicate common strategies to evade copyright infringement detectors… See the full description on the dataset page: https://huggingface.co/datasets/ml-ryanlee/free-music-archive-retrieval.audioaudio-classification1K<n<10K1 likes568 downloads1y agoHugging Face21cmu-mlsp /hubert_layer9-librispeech-asr100h Dataset Card for "hubert_layer9-librispeech-asr100h" More Information needed audio10K<n<100K0 likes547 downloads3y agoHugging Face22nimapourjafar /mm_mls_englishaudio1M<n<10M0 likes487 downloads2y agoHugging Face23cmu-mlsp /wavlm-large_layer21-librispeech-asr100h Dataset Card for "wavlm-large_layer21-librispeech-asr100h" More Information needed audio10K<n<100K0 likes451 downloads3y agoHugging Face24ky552 /ML2021_ASR_ST Dataset Card for "ML2021_ASR_ST" This dataset contains the audio recordings, the transcriptions, and the English translation of the transcriptions of the Machine Learning Course in 2021 at National Taiwan Univeristy. This can be used for domain-specific and code-switching ASR/Speech-to-text translation. If you find this dataset useful, please consider to cite the following paper: @inproceedings{yang2024investigating, title={Investigating zero-shot generalizability on… See the full description on the dataset page: https://huggingface.co/datasets/ky552/ML2021_ASR_ST.audio10K<n<100K6 likes440 downloads2y agoHugging Face25leonhard-behr /voxpopuli-mls-de-descriptionsgated Natural Language Voice Descriptions of the VoxPopuli and MLS German Datasets German read and parliamentary speech paired with its transcript, acoustic measurements, discrete German descriptor tags, and a free-text German description of the speaker's voice and recording conditions. The dataset is intended for training description-conditioned TTS models such as Parler-TTS. The data was built as part of research work. It is a random subset of the pooled German portions of VoxPopuli… See the full description on the dataset page: https://huggingface.co/datasets/leonhard-behr/voxpopuli-mls-de-descriptions.audiotext-to-speech100K<n<1M0 likes414 downloads23h agoHugging Face26thennal /indic_tts_ml Indic TTS Malayalam Speech Corpus The Malayalam subset of Indic TTS Corpus, taken from this Kaggle database. The corpus contains one male and one female speaker, with a 2:1 ratio of samples due to missing files for the female speaker. The license is given in the repository. audiotext-to-speech1K<n<10K6 likes343 downloads4y agoHugging Face27remynd /cv_mls_psfb_fs0_24Data used to reproduce all of the experiments in the paper ; https://ieeexplore.ieee.org/document/10720758/ audio100K<n<1M0 likes332 downloads1y agoHugging Face28Evan-Lin /metric-mamba-ml2021-hungyi-corpus Dataset Card for "metric-mamba-ml2021-hungyi-corpus" More Information needed audio10K<n<100K0 likes313 downloads2y agoHugging Face29philgzl /mls-hq-urgent-track1 Multilingual LibriSpeech HQ (MLS-HQ) This is a mirror of the Multilingual LibriSpeech HQ (MLS-HQ) data used in URGENT 2025 Track 1. The original files were converted from FLAC to Opus to reduce the size and accelerate streaming. Sampling rate: 48 kHz (resampled from 44.1 kHz to support Opus format) Channels: 1 Format: Opus Splits: spanish: 150 hours, 36031 utterances german: 150 hours, 35890 utterances french: 150 hours, 36078 utterances License: CC0 1.0 Source:… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/mls-hq-urgent-track1.audio100K<n<1M0 likes310 downloads4mo agoHugging Face30remynd /cv_mls_psfb_zero_syntheticData used to reproduce all of the experiments in the paper ; https://ieeexplore.ieee.org/document/10720758/ audio10K<n<100K0 likes297 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.