rasyosef/amharic-sentences-corpus
Amharic Sentences Corpus This dataset is compiled by Yimam et al. (2021) at LT Group, University of Hamburg, Germany. It comprises a collection of 6.4 million Amharic sentences intended for use in language model pretraining. Source GitHub https://github.com/uhh-lt/ethiopicmodels Dataset: https://data.mendeley.com/datasets/dtywyf3sth/1 Paper: https://www.mdpi.com/1999-5903/13/11/275 For citing this dataset, please use the following: @Article{fi13110275, AUTHOR =… See the full description on the dataset page: https://huggingface.co/datasets/rasyosef/amharic-sentences-corpus.
3138
