CoolFace
Datasetpublic

rasyosef/amharic-sentences-corpus

Amharic Sentences Corpus This dataset is compiled by Yimam et al. (2021) at LT Group, University of Hamburg, Germany. It comprises a collection of 6.4 million Amharic sentences intended for use in language model pretraining. Source GitHub https://github.com/uhh-lt/ethiopicmodels Dataset: https://data.mendeley.com/datasets/dtywyf3sth/1 Paper: https://www.mdpi.com/1999-5903/13/11/275 For citing this dataset, please use the following: @Article{fi13110275, AUTHOR =… See the full description on the dataset page: https://huggingface.co/datasets/rasyosef/amharic-sentences-corpus.

sourceHugging Faceupdated 2y agoView on Hugging Face
3likes131downloads
Dataset Card

Amharic Sentences Corpus

This dataset is compiled by Yimam et al. (2021) at LT Group, University of Hamburg, Germany. It comprises a collection of 6.4 million Amharic sentences intended for use in language model pretraining.

Source

  • GitHub https://github.com/uhh-lt/ethiopicmodels
  • Dataset: https://data.mendeley.com/datasets/dtywyf3sth/1
  • Paper: https://www.mdpi.com/1999-5903/13/11/275

For citing this dataset, please use the following:

@Article{fi13110275,
AUTHOR = {Yimam, Seid Muhie and Ayele, Abinew Ali and Venkatesh, Gopalakrishnan and Gashaw, Ibrahim and Biemann, Chris},
TITLE = {Introducing Various Semantic Models for Amharic: Experimentation and Evaluation with Multiple Tasks and Datasets},
JOURNAL = {Future Internet},
VOLUME = {13},
YEAR = {2021},
NUMBER = {11},
ARTICLE-NUMBER = {275},
URL = {https://www.mdpi.com/1999-5903/13/11/275},
ISSN = {1999-5903},
DOI = {10.3390/fi13110275}
}