datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
linguawave-competition
LinguaWave — Language Identification Competition
Pelatnas IOAI 2026 | Task 2 of 3
Identify the language of a 10-second speech clip from 8 languages. Compete to achieve the highest Macro F1-score on the test set.
Task
Input: .wav audio file (10 seconds, 16 kHz mono)Output: Language code from {id, ms, vi, th, en, zh, ar, fr}Metric: Macro F1-score
Languages
Code
Language
Region
id
Indonesian
Southeast Asia
ms
Malay
Southeast Asia
vi… See the full description on the dataset page: https://huggingface.co/datasets/fassabilf/linguawave-competition.Lingala-Speech-Dataset
Lingala Dataset Metadata
Field
Value
📜 License
CC BY-NC-ND 4.0
🎯 Task Categories
Automatic Speech Recognition
🌍 Language
Lingala (ln)
🏷️ Tags
Audio, Speech, Speech Recognition, ML, Machine, Machine Learning, Lingala
📦 Size Category
n < 1K
