datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
quranic-asr-cloud-rawdata
Quranic ASR Provider Benchmark Results
Professional benchmark artifacts for comparing commercial and official ASR providers on the Quranic ASR benchmark hosted at Quran-Lab/quranic-asr-benchmark.
This repository contains metadata, normalized result tables, raw provider responses, unchanged run scripts, scoring outputs, Tarteel streaming probes, and reports. It does not duplicate the source audio.
What Is Included
Area
Path
Purpose
Benchmark split… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quranic-asr-cloud-rawdata.nigerian-asr-dataset
NaijaVox ASR Dataset
An ASR dataset for four Nigerian languages — Hausa, Igbo, Yoruba, and
Nigerian Pidgin — sourced from Google's
Waxal corpus.
Built to train and evaluate automatic speech recognition models for
Nigerian languages.
Configs
Config
Language
Train
Validation
Test
Total
ha
Hausa
1655
296
20
1971
ig
Igbo
1604
287
20
1911
yo
Yoruba
2192
392
27
2611
pcm
Nigerian Pidgin
1674
299
20
1993
Splits: train 84% / validation 15% / test 1%… See the full description on the dataset page: https://huggingface.co/datasets/amn-raw/nigerian-asr-dataset.
