CoolFace
Datasetpublic

yordanoswuletaw/amharic-pretraining-corpus

Amharic Pretraining Corpus is a large-scale dataset (~103M) for general amharic language pretraining tasks. It consists of diverse text sources, including news articles, books, social media posts, government documents, and web content, all written in Amharic. You can load the dataset as follows from datasets import load_dataset ds = load_dataset("yordanoswuletaw/amharic-pretraining-corpus")

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
4likes266downloads
10 commits on main
21b0ca92y ago

Update README.md

yordanoswuletaw
aac37572y ago

Upload train.csv

yordanoswuletaw
22ef5ad2y ago

upload test dataset

yordanoswuletaw
bd908cb2y ago

Delete test.csv

yordanoswuletaw
52e44c72y ago

Delete eval.csv

yordanoswuletaw
cdd327d2y ago

Update README.md

yordanoswuletaw
b73c3c02y ago

Update README.md

yordanoswuletaw
8a571ee2y ago

Update README.md

yordanoswuletaw
6e9db762y ago

upload evaluation and test dataset

yordanoswuletaw
c3645a02y ago

initial commit

yordanoswuletaw