datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mls_hq_urgent_track1AT-ADD-Track1
AT-ADD Track 1
This repository hosts Track 1 of the AT-ADD All-Type Audio Deepfake Detection Challenge. It contains the released audio splits and privacy-preserving sample-level metadata for non-commercial academic research and education.
Access
This is a gated dataset. Sign in to Hugging Face, review the access agreement, complete the short access form, and click Agree and access dataset. Access is granted automatically after acceptance.
Direct repository access… See the full description on the dataset page: https://huggingface.co/datasets/xieyuankun/AT-ADD-Track1.mls-hq-urgent-track1
Multilingual LibriSpeech HQ (MLS-HQ)
This is a mirror of the Multilingual LibriSpeech HQ (MLS-HQ) data used in URGENT 2025 Track 1.
The original files were converted from FLAC to Opus to reduce the size and accelerate streaming.
Sampling rate: 48 kHz (resampled from 44.1 kHz to support Opus format)
Channels: 1
Format: Opus
Splits:
spanish: 150 hours, 36031 utterances
german: 150 hours, 35890 utterances
french: 150 hours, 36078 utterances
License: CC0 1.0
Source:… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/mls-hq-urgent-track1.vmc2026-track1-dev
vmc2026-track1-dev
Development subset of the VMC 2026 Track 1 data.
The data is organized into two configs corresponding to two subjective evaluation paradigms: absolute rating (acr) and pairwise comparison (ccr). The sample_id values are namespaced strings such as vmc2026-track1-dev-acr_489 and vmc2026-track1-dev-ccr_7233.
acr -- Absolute Category Rating
1,008 samples. Each row pairs a sample_id with one speech audio file, its released Mean Opinion Score (MOS)… See the full description on the dataset page: https://huggingface.co/datasets/urgent-challenge/vmc2026-track1-dev.vmc2026-track1-test
vmc2026-track1-test
Test subset of the VMC 2026 Track 1 data.
The data is organized into two configs corresponding to two subjective evaluation paradigms: absolute rating (acr) and pairwise comparison (ccr). The sample_id values are namespaced strings such as vmc2026-track1-test-acr_4588 and vmc2026-track1-test-ccr_3061.
acr -- Absolute Category Rating
4,032 samples. Each row pairs a sample_id with one speech audio file, its released Mean Opinion Score (MOS)… See the full description on the dataset page: https://huggingface.co/datasets/urgent-challenge/vmc2026-track1-test.urgent26_track1_leaderboard_validationresults_SE2_blind_test_dataset_v2_track1
