CoolFace
Datasetpublic

abdukuzi45/amharic-cpt-corpus-v16-balanced

📊 Dataset Token Distribution & Statistics This dataset has been cleaned and tokenized using the abdukuzi45/qwen3.5-4b-amharic-v4 tokenizer. Category (Source) Token Count Percentage Row Count 🇪🇹 Amharic 1,159,395,439 37.82% 1,244,733 💻 Code 724,855,298 23.65% 764,410 📐 Math 680,130,234 22.19% 728,000 🇬🇧 English 500,648,252 16.34% 559,054 Total 3,065,029,223 100.0% 3,296,197 Key Highlights Total Tokens: ~3.065 Billion Tokens Primary… See the full description on the dataset page: https://huggingface.co/datasets/abdukuzi45/amharic-cpt-corpus-v16-balanced.

sourceHugging Faceupdated 8d agoView on Hugging Face
0likes151downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
abdukuzi45/amharic-cpt-corpus-v16-balanced · CoolFace