mariklolik/AraToken-Qwen3-0.6B-LEP
AraToken-Qwen3-0.6B-LEP
Qwen3-0.6B-Base adapted to Arabic with the Language Extension Pipeline (LEP). 130,890 AraToken pieces were added to the vocabulary, with the new embedding rows initialized as the mean of their Qwen3 sub-token embeddings. The model was then trained on 500M tokens of FineWeb2-HQ Arabic (5,086 steps, 4 × H100). The original embedding rows were frozen by gradient masking, and only the new rows and transformer layers 24–27 were trained.
Paper: AraToken: Optimizing Arabic Tokenization with Normalization Pipeline and Language Extension for Qwen3. Code: https://github.com/mariklolik/Aratoken.
BPC is bits per character on the first 1,500 documents of the held-out test split of mariklolik/AraToken-FineWeb2-HQ-ar. It compares models with different vocabularies on the same characters.
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "mariklolik/AraToken-Qwen3-0.6B-LEP"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype="bfloat16")The tokenizer carries the AraToken normalizer: NFKC, Tatweel removal, Western digits, Latin punctuation and diacritics removal. Raw Arabic text can therefore be passed to it directly.
Citation
@article{kashirskiy2025aratoken,
title = {AraToken: Optimizing Arabic Tokenization with Normalization Pipeline and Language Extension for Qwen3},
author = {Kashirskiy, Mark and Lipinski, Artiom and Makarov, Ilya},
journal = {arXiv preprint arXiv:2512.18399},
year = {2025}
}