datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MLAAD
Introduction
Welcome to MLAAD: The Multi-Language Audio Anti-Spoofing Dataset -- a dataset to train, test and evaluate audio deepfake detection. See
the paper for more information.
License
MLAAD is published strictly for non-commercial academic research use, under the CC-BY-NC 4.0 license. Commercial use is not permitted.
Bibtex
If you use this dataset, please consider citing it as follows.
@article{muller2024mlaad,
title={MLAAD: The… See the full description on the dataset page: https://huggingface.co/datasets/mueller91/MLAAD.MLAAD-tiny
Welcome to MLAAD-tiny
MLAAD-tiny is a very small subset of the full MLAAD dataset, designed for education, prototyping, and debugging.
Many teaching environments (e.g. Colab, Kaggle, university notebooks -- se this notebook for example) impose strict storage limits, which makes large-scale audio deepfake datasets impractical to use. To address this, we provide MLAAD-tiny, a compact yet representative version of MLAAD.
Download
git lfs install
git clone… See the full description on the dataset page: https://huggingface.co/datasets/mueller91/MLAAD-tiny.MCL-MLAAD
Introduction
MCL-MLAAD is the first multilingual benchmark for speech deepfake source tracing. It spans mono- and cross-lingual protocols, includes DSP and SSL baselines, studies language-specific fine-tuning for cross-lingual generalization, and tests robustness to unseen languages/speakers. See arXiv:2508.04143.
Download the Dataset
Install the datasets package:
pip install datasets
Log in with your Hugging Face account:
huggingface-cli login
Load the dataset in… See the full description on the dataset page: https://huggingface.co/datasets/xxuan-speech/MCL-MLAAD.DFADD_MLAAD_DiffSSD_VoxCeleb2MLAAD-tiny
Welcome to MLAAD-tiny
MLAAD-tiny is a very small subset of the full MLAAD dataset, designed for education, prototyping, and debugging.
Many teaching environments (e.g. Colab, Kaggle, university notebooks) impose strict storage limits, which makes large-scale audio deepfake datasets impractical to use. To address this, we provide MLAAD-tiny, a compact yet representative version of MLAAD.
Download
git lfs install
git clone… See the full description on the dataset page: https://huggingface.co/datasets/Saisaket25/MLAAD-tiny.v14-MLAAD-Fake-part_01v14-MLAAD-Fake-part_05v14-MLAAD-Fake-part_02v14-MLAAD-Fake-part_03v14-MLAAD-Fake-part_04MLAAD_Protocols_for_Source-Speaker_Disentanglement_Researchv14-MLAAD-Fake-part_10v14-MLAAD-Fake-part_07v14-MLAAD-Fake-part_08v14-MLAAD-Fake-part_06v14-MLAAD-Fake-part_09MLAAD_Audit
MLAAD — SSA Acoustic Feature Audit
Moonscape Software | Synthetic Speech Atlas
Research audit contribution to the MLAAD dataset team
Overview
This repository contains acoustic feature measurements extracted from the
MLAAD (Multilingual Audio Anti-Spoofing Dataset) corpus by the Moonscape
Synthetic Speech Atlas (SSA) pipeline.
298,000 rows. 152 columns. One row per MLAAD clip.
No audio files are included. Each row contains classical signal processing
and biomechanical… See the full description on the dataset page: https://huggingface.co/datasets/moonscape-software/MLAAD_Audit.MLAAD_v8mlaad-2000-audioMLAAD-chunked
