datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NADI2026_subtask1.1_Robust_ASR
Dataset Card for "NADI2026_subtask1.1_Robust_ASR"
More Information needed
NADI2026_subtask2_MixedASR
Dataset Card for "NADI2026_subtask2_MixedASR"
More Information needed
nadi_2026_subtask_4_testNADI2026_subtask1.1_Robust_ASR_test
Dataset Card for "NADI2026_subtask1.1_Robust_ASR_test"
More Information needed
NADI2026_subtask2_adi_testNADI2026_subtask1.3_codeswitched_asrNADI2026_subtask1.2_MixedASR_test
Dataset Card for "NADI2026_subtask2_MixedASR_test"
More Information needed
nadi26_subtask_5_SLU_test
NADI 2026 Official Test Set
Official test set for NADI 2026 Subtask 5: Spoken Language Understanding.
Repository:
https://huggingface.co/datasets/UBC-NLP/nadi26_subtask_5_SLU_test
The dataset contains the clean test split with 989 examples.
Each example includes:
audio
id
The audio is embedded directly in Parquet.
NADI2025_subtask1_SLIDTo participate in the NADI 2025 Spoken Dialect ID challenge please make sure you have visited the main NADI 2025 page link, and sign the participation form on CodaBench. Ensure your email + contact information matches your Huggingface Access request email.
This is the `adaptation' split for the NADI 2025 Spoken Dialect ID task. As an adaptation split, the idea is to use an existing external dataset (e.g. ADI-17) for the main training, and then use this split for fine-tuning ('train') and… See the full description on the dataset page: https://huggingface.co/datasets/UBC-NLP/NADI2025_subtask1_SLID.
