dagaare
Datasets
All datasets matching “dagaare”dagaare_synth_trainThe dagaare_dict_guided_train_9k.tsv presents the synthesized data for the dictionary-guided training set, which is designed on the observed dictionary "A dictionary and grammatical sketch of Dagaare" by Ali, Grimm, and Bodomo (2021).
Dictionary terms were selected as target words in the curation of the dataset.
dagaare_synth_source_targetThe dagaareDictTrain.tsv was generated using the Machine Translation from One Book (MTOB) technique using "A dictionary and grammatical sketch of Dagaare" by Ali, Grimm, and Bodomo (2021) for LLM context.
dagaare-speech-data
Dagaare Speech Data (Pooled)
A ~104.3-hour Dagaare (Dagara) speech corpus, drawn from a single source
(WAXAL) and filtered to only genuinely transcribed audio. Part of the
AfroNet multi-language TTS data
effort.
Source
WAXAL (google/WaxalNLP),
dga_asr config — crowdsourced, image-prompted speech (a shared collection pipeline
also used for Dagbani, Ikposo, and Akan's aka_asr in this collection). 18,859
clips, 104.3h, source = waxal.
A known upstream bug, verified… See the full description on the dataset page: https://huggingface.co/datasets/Professor/dagaare-speech-data.
