intronhealth/afrispeech-200
AFRISPEECH-200 is a 200hr Pan-African speech corpus for clinical and general domain English accented ASR; a dataset with 120 African accents from 13 countries and 2,463 unique African speakers. Our goal is to raise awareness for and advance Pan-African English ASR research, especially for the clinical domain.
add dev and test transcripts
updated transcript
Update README.md
unlock test set audios
fix absent split meta
fix absent splits
fix absent splits
fix split indexing for single-accent configs
fix -all- config
fix audio shard loop
fix audio shard loop
fix accent keys
fix accent keys
fix shards
remove duplicates
update readme with config for individual accents
fix list bug
fix list bug
enable single accent download
add audio_id field
update readme with all config
remove dev transcripts
mute search logs
add documentation on download size and streaming
add documentation on download size and streaming
add colab link
add documentation on download size and streaming
add tqdm to audio search
add config logging
fix constants
add configs for smaller datasets
fix metadata reading
fix audio key in metadata
remove default config name
fix data urls
add train csv
upload tar files and csvs
initial commit
