amanuelbyte/Amharic_dataset
Amharic Articles Dataset Dataset Description This dataset comprises news articles in Amharic. Dataset Details Language: Amharic Time Coverage: Primarily December 2020 and January 2021 (Ethiopian Calendar: Tahisas 2013 and Tir 2013) Data Format: Each entry is a text line within a .txt text corpus. Potential Uses This dataset can be used for various purposes: Pretraining Amharic Language Models: The dataset can be used for… See the full description on the dataset page: https://huggingface.co/datasets/amanuelbyte/Amharic_dataset.
053
Amharic Articles Dataset
Dataset Description
This dataset comprises news articles in Amharic.
Dataset Details
- Language: Amharic
- Time Coverage: Primarily December 2020 and January 2021 (Ethiopian Calendar: Tahisas 2013 and Tir 2013)
- Data Format: Each entry is a text line within a
.txttext corpus.
Potential Uses
This dataset can be used for various purposes:
- Pretraining Amharic Language Models: The dataset can be used for pretraining new Amharic language models from scratch.
- Continued Pretraining: It is also suitable for continued pretraining of existing Amharic language models to adapt them to specific domains or update their knowledge.
- Text Summarization: Generating concise summaries from long news articles.
- Natural Language Processing (NLP) Research: Conducting research on Amharic text analysis, word embeddings, and other NLP tasks.
- Historical and Social Science Research: Analyzing major events and trends reported during a specific period in Ethiopia.
- Machine Translation: Training models for translation from Amharic to other languages.
Limitations and Considerations
- Data Cleaning: As the dataset is in raw text format, it may require cleaning of formatting issues (e.g., inconsistent lines, incomplete information) before use for analysis.
- Licensing: Explicit licensing information for using the dataset is not provided. Users should verify copyright and usage guidelines when acquiring the data.
- Dataset Size: The provided dataset size is small and may not be sufficient for training large machine learning models.
contact
geaman73@gmail.com
