CoolFace
Datasetpublic

intronhealth/afrispeech-200

AFRISPEECH-200 is a 200hr Pan-African speech corpus for clinical and general domain English accented ASR; a dataset with 120 African accents from 13 countries and 2,463 unique African speakers. Our goal is to raise awareness for and advance Pan-African English ASR research, especially for the clinical domain.

sourceHugging Facecc-by-nc-sa-4.0updated 3y agoView on Hugging Face
40likes2kdownloads
38 commits on main
b538c6e3y ago

add dev and test transcripts

Owos
704a0d63y ago

updated transcript

Owos
8ccec763y ago

Update README.md

tobiolatunji
4701e103y ago

unlock test set audios

tobiolatunji
9d61f374y ago

fix absent split meta

tobiolatunji
472576b4y ago

fix absent splits

tobiolatunji
31583b54y ago

fix absent splits

tobiolatunji
7dd0ea44y ago

fix split indexing for single-accent configs

polinaeterna
e4f7b4f4y ago

fix -all- config

tobiolatunji
1042eff4y ago

fix audio shard loop

tobiolatunji
931cf654y ago

fix audio shard loop

tobiolatunji
7f89ccd4y ago

fix accent keys

tobiolatunji
57c20fd4y ago

fix accent keys

tobiolatunji
7c0e4064y ago

fix shards

tobiolatunji
8a9e6434y ago

remove duplicates

tobiolatunji
50745c84y ago

update readme with config for individual accents

tobiolatunji
0b720c84y ago

fix list bug

tobiolatunji
6e27c314y ago

fix list bug

tobiolatunji
44114024y ago

enable single accent download

tobiolatunji
48be34f4y ago

add audio_id field

tobiolatunji
db5d2e34y ago

update readme with all config

tobiolatunji
e9dead84y ago

remove dev transcripts

tobiolatunji
9c93bb14y ago

mute search logs

tobiolatunji
caf29844y ago

add documentation on download size and streaming

tobiolatunji
89577314y ago

add documentation on download size and streaming

tobiolatunji
697b2fe4y ago

add colab link

tobiolatunji
b3010f54y ago

add documentation on download size and streaming

tobiolatunji
d0f6eb54y ago

add tqdm to audio search

tobiolatunji
b179fcd4y ago

add config logging

tobiolatunji
b74379f4y ago

fix constants

tobiolatunji
75fceb64y ago

add configs for smaller datasets

tobiolatunji
27520da4y ago

fix metadata reading

polinaeterna
74224494y ago

fix audio key in metadata

polinaeterna
ca906e04y ago

remove default config name

polinaeterna
1eecf834y ago

fix data urls

polinaeterna
e594a174y ago

add train csv

tobiolatunji
a239d5a4y ago

upload tar files and csvs

tobiolatunji
41181f74y ago

initial commit

tobiolatunji