rasyosef/amharic-sentences-corpus
Amharic Sentences Corpus This dataset is compiled by Yimam et al. (2021) at LT Group, University of Hamburg, Germany. It comprises a collection of 6.4 million Amharic sentences intended for use in language model pretraining. Source GitHub https://github.com/uhh-lt/ethiopicmodels Dataset: https://data.mendeley.com/datasets/dtywyf3sth/1 Paper: https://www.mdpi.com/1999-5903/13/11/275 For citing this dataset, please use the following: @Article{fi13110275, AUTHOR =… See the full description on the dataset page: https://huggingface.co/datasets/rasyosef/amharic-sentences-corpus.
Amharic Sentences Corpus
This dataset is compiled by Yimam et al. (2021) at LT Group, University of Hamburg, Germany. It comprises a collection of 6.4 million Amharic sentences intended for use in language model pretraining.
Source
- GitHub https://github.com/uhh-lt/ethiopicmodels
- Dataset: https://data.mendeley.com/datasets/dtywyf3sth/1
- Paper: https://www.mdpi.com/1999-5903/13/11/275
For citing this dataset, please use the following:
@Article{fi13110275,
AUTHOR = {Yimam, Seid Muhie and Ayele, Abinew Ali and Venkatesh, Gopalakrishnan and Gashaw, Ibrahim and Biemann, Chris},
TITLE = {Introducing Various Semantic Models for Amharic: Experimentation and Evaluation with Multiple Tasks and Datasets},
JOURNAL = {Future Internet},
VOLUME = {13},
YEAR = {2021},
NUMBER = {11},
ARTICLE-NUMBER = {275},
URL = {https://www.mdpi.com/1999-5903/13/11/275},
ISSN = {1999-5903},
DOI = {10.3390/fi13110275}
}
