CoolFace
Modelpublic

NaolBM/Africa-BBPE

sourceHugging Facemitupdated 7mo agoView on Hugging Face
0likes
Model Card

Africa-BBPE Tokenizer

๐Ÿ“‹ Model Card

Model Overview

  • โ€”Model Name: NaolBM/Africa-BBPE
  • โ€”Model Type: Byte-Level BPE Tokenizer
  • โ€”Base Model: Trained on NaolBM/african-corpus
  • โ€”Vocabulary Size: 50000
  • โ€”Languages Supported: Amharic, Swahili, Hausa, Oromo, Yoruba, Tigrinya, English

๐ŸŽฏ Performance Benchmarks

Tokenizer Battle Results

Comparison against Gemma-3 and Qwen3 tokenizers:

LanguageAfrica-BBPEGemma-3Qwen-3Winner
Amharic41518๐Ÿ‡ช๐Ÿ‡น Africa-BBPE
Swahili81012๐Ÿ‡ฐ๐Ÿ‡ช Africa-BBPE
Hausa111212๐Ÿ‡ณ๐Ÿ‡ฌ Africa-BBPE
Oromo51110๐Ÿ‡ช๐Ÿ‡น Africa-BBPE
Yoruba988๐Ÿ‡ณ๐Ÿ‡ฌ Gemma-3
Tigrinya51321๐Ÿ‡ช๐Ÿ‡ท Africa-BBPE
English743๐Ÿ‡ฌ๐Ÿ‡ง Qwen-3
Code-switching101721๐ŸŒ Africa-BBPE

Summary Statistics

MetricAfrica-BBPEGemma-3Qwen-3
๐Ÿ† Wins611
๐Ÿ“Š Total Tokens5990105
โšก Avg Tokens/Sample7.3811.2513.13

Language Family Performance

Language FamilyAfrica-BBPEGemma-3Qwen-3
Semitic (Ge'ez)4.514.019.5
Cushitic5.011.010.0
Bantu8.010.012.0
Chadic11.012.012.0
Benue-Congo9.08.08.0
Germanic7.04.03.0
Code-switching10.017.021.0

๐ŸŒ Language Support

LanguageCodeScriptTokenization Efficiency
AmharicamGe'ezโญโญโญโญโญ (4 tokens avg)
TigrinyatiGe'ezโญโญโญโญโญ (5 tokens avg)
OromoomLatinโญโญโญโญโญ (5 tokens avg)
SwahiliswLatinโญโญโญโญ (8 tokens avg)
HausahaLatinโญโญโญ (11 tokens avg)
YorubayoLatinโญโญโญ (9 tokens avg)
EnglishenLatinโญโญ (7 tokens avg)
Code-switchingMixedMixedโญโญโญโญโญ (10 tokens avg)

๐Ÿš€ Usage

python
from transformers import AutoTokenizer

# Load tokenizer
tokenizer = AutoTokenizer.from_pretrained("NaolBM/Africa-BBPE")

# Example usage
text = "แŠ แˆ›แˆญแŠ› แ‰‹แŠ•แ‰‹ แ‰ แŠขแ‰ตแ‹ฎแŒตแ‹ซ"
tokens = tokenizer.tokenize(text)
ids = tokenizer.encode(text)

print(f"Tokens: {tokens}")
print(f"Token IDs: {ids}")
print(f"Number of tokens: {len(ids)}")

๐Ÿ“Š Training Data

  • โ€”Dataset: NaolBM/african-corpus
  • โ€”Total Rows: 35,344,339
  • โ€”Languages: Amharic, Swahili, Hausa, Oromo, Yoruba, Tigrinya, English

Dataset Composition

LanguageRowsPercentage
Swahili14,125,92539.97%
Amharic10,815,25530.60%
Hausa7,144,07720.21%
English2,119,7196.00%
Oromo881,4502.49%
Yoruba245,8370.70%
Tigrinya12,0760.03%

๐Ÿ” Key Strengths

  • โ€”โœ… 6x more efficient than Gemma-3 on Amharic (4 vs 15 tokens)
  • โ€”โœ… 4x more efficient than Qwen-3 on Tigrinya (5 vs 21 tokens)
  • โ€”โœ… 2x more efficient on code-switched text (10 vs 17-21 tokens)
  • โ€”โœ… Best-in-class for 6/8 language categories
  • โ€”โœ… Optimized for Ge'ez script (Amharic, Tigrinya)
  • โ€”โœ… Excellent for Cushitic languages (Oromo)

๐Ÿ“ˆ Efficiency Gains

Compared to Gemma-3:

  • โ€”73% fewer tokens for Amharic
  • โ€”62% fewer tokens for Tigrinya
  • โ€”55% fewer tokens for Oromo
  • โ€”41% fewer tokens overall

Compared to Qwen-3:

  • โ€”78% fewer tokens for Tigrinya
  • โ€”72% fewer tokens for code-switching
  • โ€”52% fewer tokens for Amharic
  • โ€”44% fewer tokens overall

๐Ÿ“œ License

MIT

๐Ÿ™ Acknowledgments

  • โ€”Trained on NaolBM/african-corpus
  • โ€”Built with Hugging Face Tokenizers library
  • โ€”All original dataset contributors

๐Ÿ”— Links