CoolFace
Datasetpublic

rasyosef/amharic-sentences-corpus

Amharic Sentences Corpus This dataset is compiled by Yimam et al. (2021) at LT Group, University of Hamburg, Germany. It comprises a collection of 6.4 million Amharic sentences intended for use in language model pretraining. Source GitHub https://github.com/uhh-lt/ethiopicmodels Dataset: https://data.mendeley.com/datasets/dtywyf3sth/1 Paper: https://www.mdpi.com/1999-5903/13/11/275 For citing this dataset, please use the following: @Article{fi13110275, AUTHOR =… See the full description on the dataset page: https://huggingface.co/datasets/rasyosef/amharic-sentences-corpus.

sourceHugging Faceupdated 2y agoView on Hugging Face
3likes138downloads

rasyosef/amharic-sentences-corpus · main · files are served by the source, never re-hosted here