CoolFace
Datasetpublic

abdukuzi45/amharic-cpt-corpus-v16-balanced

πŸ“Š Dataset Token Distribution & Statistics This dataset has been cleaned and tokenized using the abdukuzi45/qwen3.5-4b-amharic-v4 tokenizer. Category (Source) Token Count Percentage Row Count πŸ‡ͺπŸ‡Ή Amharic 1,159,395,439 37.82% 1,244,733 πŸ’» Code 724,855,298 23.65% 764,410 πŸ“ Math 680,130,234 22.19% 728,000 πŸ‡¬πŸ‡§ English 500,648,252 16.34% 559,054 Total 3,065,029,223 100.0% 3,296,197 Key Highlights Total Tokens: ~3.065 Billion Tokens Primary… See the full description on the dataset page: https://huggingface.co/datasets/abdukuzi45/amharic-cpt-corpus-v16-balanced.

sourceHugging Faceupdated 8d agoView on Hugging Face
0likes151downloads
Dataset Card

πŸ“Š Dataset Token Distribution & Statistics

This dataset has been cleaned and tokenized using the `abdukuzi45/qwen3.5-4b-amharic-v4` tokenizer.

Category (Source)Token CountPercentageRow Count
πŸ‡ͺπŸ‡Ή Amharic1,159,395,43937.82%1,244,733
πŸ’» Code724,855,29823.65%764,410
πŸ“ Math680,130,23422.19%728,000
πŸ‡¬πŸ‡§ English500,648,25216.34%559,054
Total3,065,029,223100.0%3,296,197

Key Highlights

  • β€”Total Tokens: ~3.065 Billion Tokens
  • β€”Primary Language: Amharic (~1.16 Billion Tokens)
  • β€”Quality Control: 100% filtered from auto-generated C++ bindings, IL2CPP hashes, logic gate netlists, and minified code snippets.