datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
voiceguard-competition
VoiceGuard — Deepfake Audio Detection Competition
Pelatnas IOAI 2026 | Task 3 of 3
Detect whether a 4-second audio clip is real human speech or AI-generated (TTS/deepfake). Submit probability scores — AUROC is the metric.
Task
Input: .wav audio file (4 seconds, 16 kHz mono)Output: score — probability (0–1) that the audio is fakeMetric: AUROC (Area Under ROC Curve)
Dataset
Split
Real
Fake
Total
Train
2,874
2,874
5,748
Test
627
627
1,254… See the full description on the dataset page: https://huggingface.co/datasets/fassabilf/voiceguard-competition.linguawave-competition
LinguaWave — Language Identification Competition
Pelatnas IOAI 2026 | Task 2 of 3
Identify the language of a 10-second speech clip from 8 languages. Compete to achieve the highest Macro F1-score on the test set.
Task
Input: .wav audio file (10 seconds, 16 kHz mono)Output: Language code from {id, ms, vi, th, en, zh, ar, fr}Metric: Macro F1-score
Languages
Code
Language
Region
id
Indonesian
Southeast Asia
ms
Malay
Southeast Asia
vi… See the full description on the dataset page: https://huggingface.co/datasets/fassabilf/linguawave-competition.
