CoolFace
Datasetpublic

SpeechAntiSpoofingBenchmarks/ASVspoof5

ASVspoof 5 (track 1, eval) Benchmark-ready packaging of the Track 1 (spoofing / deepfake detection) evaluation partition of the ASVspoof 5 challenge, for speech anti-spoofing and synthetic / deepfake voice detection. Overview Track 1 is binary classification: bonafide (genuine human speech) vs. spoof (synthetic / converted speech). This packaging contains the full track_1 evaluation set. The original challenge is at https://www.asvspoof.org/. License &… See the full description on the dataset page: https://huggingface.co/datasets/SpeechAntiSpoofingBenchmarks/ASVspoof5.

sourceHugging Faceodc-byupdated 3mo agoView on Hugging Face
0likes2.7kdownloads
Dataset Card

ASVspoof 5 (track 1, eval)

Benchmark-ready packaging of the Track 1 (spoofing / deepfake detection) evaluation partition of the ASVspoof 5 challenge, for speech anti-spoofing and synthetic / deepfake voice detection.

Overview

Track 1 is binary classification: bonafide (genuine human speech) vs. spoof (synthetic / converted speech). This packaging contains the full track_1 evaluation set. The original challenge is at https://www.asvspoof.org/.

License & redistribution

Redistributed under the Open Data Commons Attribution License (ODC-By) v1.0. See LICENSE.txt. Labels and the evaluation protocol are unmodified; audio is the original 16 kHz mono FLAC, embedded bit-exactly (no re-encode — a full decode probe of all 680,774 clips passed cleanly).

Schema

ColumnTypeDescription
pathstring<utterance_id>.flac, unique
audioAudio(16000)16 kHz mono FLAC
labelClassLabel"bonafide" (0) / "spoof" (1)
notesstringJSON: utterance_id, speaker_id, gender, codec, codec_id, source_id, attack_condition, attack_id

notes example:

json
{"utterance_id": "E_0009538969", "speaker_id": "E_1607", "gender": "M", "codec": "C05", "codec_id": "2", "source_id": "E_0009486171", "attack_condition": "AC1", "attack_id": "A26"}

Quick Start

python
from datasets import load_dataset

ds = load_dataset("SpeechAntiSpoofingBenchmarks/ASVspoof5", split="test")
print(ds[0])

Stats

StatValue
Total trials680,774
Bonafide138,688
Spoof542,086

Source provenance

  • —Original challenge: https://www.asvspoof.org/
  • —Evaluation protocol: ASVspoof5.eval.track_1.tsv

Evaluation

For evaluation instructions and submission format, see `submissions/README.md`.

Citation

bibtex
@inproceedings{wang2024asvspoof5,
  title     = {{ASVspoof 5: Crowdsourced Speech Data, Deepfakes, and Adversarial Attacks at Scale}},
  author    = {Wang, Xin and Delgado, H{\'e}ctor and Tak, Hemlata and others},
  year      = {2024},
  booktitle = {ASVspoof Workshop 2024},
}

Maintainer

Maintained by Kirill Borodin (SpeechAntiSpoofingBenchmarks).

  • —Email: ~~k.n.borodin@mtuci.ru~~ (deprecated — use kborodin.research@gmail.com)
  • —Telegram: @korallll_ai