yordanoswuletaw/amharic-pretraining-corpus
Amharic Pretraining Corpus is a large-scale dataset (~103M) for general amharic language pretraining tasks. It consists of diverse text sources, including news articles, books, social media posts, government documents, and web content, all written in Amharic. You can load the dataset as follows from datasets import load_dataset ds = load_dataset("yordanoswuletaw/amharic-pretraining-corpus")
4293
This repository belongs to yordanoswuletaw on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
amharic-pretraining-corpus
public
apache-2.0
no
yordanoswuletaw
