datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dagaare_synth_trainThe dagaare_dict_guided_train_9k.tsv presents the synthesized data for the dictionary-guided training set, which is designed on the observed dictionary "A dictionary and grammatical sketch of Dagaare" by Ali, Grimm, and Bodomo (2021).
Dictionary terms were selected as target words in the curation of the dataset.
dagaare_synth_source_targetThe dagaareDictTrain.tsv was generated using the Machine Translation from One Book (MTOB) technique using "A dictionary and grammatical sketch of Dagaare" by Ali, Grimm, and Bodomo (2021) for LLM context.
