mariklolik/AraToken-Qwen3-1.7B-CPT
0295
AraToken-Qwen3-1.7B-CPT
The continued-pretraining baseline at 1.7B parameters. It uses the original tokenizer and the same data, trainable layers (24–27) and number of steps (2,000) as mariklolik/AraToken-Qwen3-1.7B-LEP.
Paper: AraToken: Optimizing Arabic Tokenization with Normalization Pipeline and Language Extension for Qwen3. Code: https://github.com/mariklolik/Aratoken.
BPC is bits per character on the first 1,500 documents of the held-out test split of mariklolik/AraToken-FineWeb2-HQ-ar. It compares models with different vocabularies on the same characters.
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "mariklolik/AraToken-Qwen3-1.7B-CPT"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype="bfloat16")The tokenizer carries the AraToken normalizer: NFKC, Tatweel removal, Western digits, Latin punctuation and diacritics removal. Raw Arabic text can therefore be passed to it directly.
Citation
@article{kashirskiy2025aratoken,
title = {AraToken: Optimizing Arabic Tokenization with Normalization Pipeline and Language Extension for Qwen3},
author = {Kashirskiy, Mark and Lipinski, Artiom and Makarov, Ilya},
journal = {arXiv preprint arXiv:2512.18399},
year = {2025}
}