CoolFace
Datasetpublic

ghanaopenai/ghana-speech-eval

ghana-speech-eval ⚠️ Evaluation only. This is a held-out benchmark for measuring ASR quality on Ghanaian languages. Please do not train or fine-tune on it — doing so contaminates the benchmark and makes reported scores meaningless. Every subset ships a single eval split. Multi-source ASR evaluation benchmark for Ghanaian languages. All subsets are 16 kHz mono with standardised schema (audio, text, language, country, length, iso, subset). Source groups… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-speech-eval.

sourceHugging Faceupdated 29d agoView on Hugging Face
0likes581downloads
Dataset Card

ghana-speech-eval

⚠️ Evaluation only. This is a held-out benchmark for measuring ASR quality on Ghanaian languages. Please do not train or fine-tune on it — doing so contaminates the benchmark and makes reported scores meaningless. Every subset ships a single eval split.

Multi-source ASR evaluation benchmark for Ghanaian languages. All subsets are 16 kHz mono with standardised schema (audio, text, language, country, length, iso, subset).

Source groups

GroupDescriptionClip lengthPer-language cap
bible_*Bible-audio aligned (ghananlpcommunity/ghana-speech)3–15 s1 000
finance_*Finance-domain multispeaker speech (ghanaopendata)0–60 s1 000
jw_*JW.org speech (AfriSpeech v1)3–15 s1 000
lds_*LDS general-conference speech (multispeaker)3–15 s1 000
unicef_*UNICEF health recordings3–60 sall available (~195–200)
waxal_*WaxalNLP ASR (google/WaxalNLP)3–30 s1 000

Twi is split by dialect

Akuapem and Asante Twi are separate languages here, twi_akuapem and twi_asante. They were previously merged into a single twi (500/500 in bible_twi_twi and finance_twi), which made every Twi score an average over two orthographies: the Khaya API's Akuapem endpoint scored 8.9 % WER on the Akuapem-heavy bible split and 88.7 % on jw, and its general-Twi endpoint the reverse. The merged configs (bible_twi_twi, finance_twi, unicef_twi, jw_twi_twi) have been removed. jw_* remains for the other nine languages — its Twi audio mixes both dialects with no way to separate them.

Neither dialect has its own ISO 639-3 code (both are twi), so twi_akuapem and twi_asante are this dataset's keys, not ISO.

Two labels are our classification rather than the source's, and are stated here so they are not mistaken for ground truth:

  • —lds_Asante_Twi — the source is labelled only "Twi (Akan)"; its orthography (mfeɛ, afoforɔ) matches the verified Asante bible text.
  • —waxal_Asante_Twi — WaxalNLP labels this config aka / "Akan" without a dialect; it is filed under Asante Twi.

Configs (65 total)

bible_* — 42 configs

ConfigISOLanguageClips
bible_Gikyode_acdacdGikyode1000
bible_Dangme_adaadaDangme1000
bible_Siwu_akpakpSiwu1000
bible_Anyin_anyanyAnyin1000
bible_Avatime_avnavnAvatime1000
bible_Bissa_bibbibBissa1000
bible_Bimoba_bimbimBimoba1000
bible_Birifor_Southern_bivbivBirifor_Southern1000
bible_Tuwuli_bovbovTuwuli1000
bible_Bassar_Ntcham_budbudBassar_Ntcham1000
bible_Buli_bwubwuBuli1000
bible_Dagbani_dagdagDagbani1000
bible_Dagaare_dgadgaDagaare1000
bible_Ewe_eweeweEwe1000
bible_Fante_fatfatFante1000
bible_Fulfulde_Maasina_ffmffmFulfulde_Maasina1000
bible_Gonja_gjngjnGonja1000
bible_Ninkare_gurgurNinkare1000
bible_Hausa_hauhauHausa1000
bible_Kabiye_kbpkbpKabiye1000
bible_Tem_kdhkdhTem1000
bible_Konni_kmakmaKonni1000
bible_Kusaal_kuskusKusaal1000
bible_Lelemi_leflefLelemi1000
bible_Sekpele_liplipSekpele1000
bible_Mampruli_mawmawMampruli1000
bible_Deg_mzwmzwDeg1000
bible_Nawuri_nawnawNawuri1000
bible_Chumburung_ncuncuChumburung1000
bible_Nkonya_nkonkoNkonya1000
bible_Ntrubo_ntrntrNtrubo1000
bible_Nzema_nzinziNzema1000
bible_Sehwi_sfwsfwSehwi1000
bible_Paasaal_sigsigPaasaal1000
bible_Sisaala_Tumulung_silsilSisaala_Tumulung1000
bible_Selee_snwsnwSelee1000
bible_Tampulma_tpmtpmTampulma1000
bible_Akuapem_Twitwi_akuapemAkuapem Twi1000
bible_Asante_Twitwi_asanteAsante Twi1000
bible_Vagla_vagvagVagla1000
bible_Konkomba_xonxonKonkomba1000
bible_Kasem_xsmxsmKasem1000

finance_* — 4 configs

ConfigISOLanguageClips
finance_fantefatFante1000
finance_gagaaGa1000
finance_Akuapem_Twitwi_akuapemAkuapem Twi1000
finance_Asante_Twitwi_asanteAsante Twi1000

jw_* — 9 configs

ConfigISOLanguageClips
jw_dangme_adaadaDangme261
jw_ahanta_ahaahaAhanta1000
jw_dagaare_dgadgaSouthern Dagaare1000
jw_ewe_eweeweÉwé1000
jw_fante_fatfatFante1000
jw_ga_gaagaaGa1000
jw_frafra_gurgurFarefare1000
jw_nzema_nzinziNzema1000
jw_sehwi_sfwsfwEsahie1000

lds_* — 2 configs

ConfigISOLanguageClips
lds_Fante_fatfatFante1000
lds_Asante_Twitwi_asanteAsante Twi1000

unicef_* — 3 configs

ConfigISOLanguageClips
unicef_dagbanidagDagbani195
unicef_eweeweÉwé199
unicef_Asante_Twitwi_asanteAsante Twi196

waxal_* — 5 configs

ConfigISOLanguageClips
waxal_Dagbani_dagdagDagbani1000
waxal_Dagaare_dgadgaDagaare1000
waxal_Ewe_eweeweEwe1000
waxal_Ikposo_kpokpoIkposo1000
waxal_Asante_Twitwi_asanteAsante Twi1000

Loading

python
from datasets import load_dataset

# every subset has a single `eval` split
ds = load_dataset("ghananlpcommunity/ghana-speech-eval", "bible_Asante_Twi", split="eval")

from datasets import get_dataset_config_names
print(get_dataset_config_names("ghananlpcommunity/ghana-speech-eval"))

Schema

ColumnTypeDescription
audioAudio16 kHz mono PCM
textstringreference transcription (verbatim)
languagestringhuman-readable language label
countrystringISO 3166-1 alpha-2, or ""
lengthfloat64clip duration in seconds
isostringlanguage key (ISO 639-3, except the two Twi dialect keys)
subsetstring<group>_<label>, matching the config name

Built with speecheval-builder. Benchmarked by nsanku-ASR.

License

CC-BY-4.0. Respect the terms of each source dataset.