abdukuzi45/amharic-cpt-corpus-v16-balanced
📊 Dataset Token Distribution & Statistics This dataset has been cleaned and tokenized using the abdukuzi45/qwen3.5-4b-amharic-v4 tokenizer. Category (Source) Token Count Percentage Row Count 🇪🇹 Amharic 1,159,395,439 37.82% 1,244,733 💻 Code 724,855,298 23.65% 764,410 📐 Math 680,130,234 22.19% 728,000 🇬🇧 English 500,648,252 16.34% 559,054 Total 3,065,029,223 100.0% 3,296,197 Key Highlights Total Tokens: ~3.065 Billion Tokens Primary… See the full description on the dataset page: https://huggingface.co/datasets/abdukuzi45/amharic-cpt-corpus-v16-balanced.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face