abdukuzi45/amharic-cpt-corpus-v16-balanced
π Dataset Token Distribution & Statistics This dataset has been cleaned and tokenized using the abdukuzi45/qwen3.5-4b-amharic-v4 tokenizer. Category (Source) Token Count Percentage Row Count πͺπΉ Amharic 1,159,395,439 37.82% 1,244,733 π» Code 724,855,298 23.65% 764,410 π Math 680,130,234 22.19% 728,000 π¬π§ English 500,648,252 16.34% 559,054 Total 3,065,029,223 100.0% 3,296,197 Key Highlights Total Tokens: ~3.065 Billion Tokens Primaryβ¦ See the full description on the dataset page: https://huggingface.co/datasets/abdukuzi45/amharic-cpt-corpus-v16-balanced.
0151
π Dataset Token Distribution & Statistics
This dataset has been cleaned and tokenized using the `abdukuzi45/qwen3.5-4b-amharic-v4` tokenizer.
Key Highlights
- Total Tokens: ~3.065 Billion Tokens
- Primary Language: Amharic (~1.16 Billion Tokens)
- Quality Control: 100% filtered from auto-generated C++ bindings, IL2CPP hashes, logic gate netlists, and minified code snippets.
