carry121/k3-32k-tokenizer
0
k3-32k tokenizer
Bilingual (Chinese / English) SentencePiece Unigram tokenizer with a 32,000-piece vocabulary, trained for the K3/KQ model family.
Usage
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("carry121/k3-32k-tokenizer")
ids = tok.encode("深度学习是机器学习的一个分支。")
# BOS (<s>, id=2) is prepended automatically (add_bos_token=True).Both fast (tokenizers backend) and slow (SentencePiece) loading work; encodings are identical.
Special tokens
IDs 0–3 are fixed and stable; the four user-defined symbols always encode to a single id.
Training
- Algorithm: SentencePiece Unigram, NFKC normalization,
character_coverage=0.9995,byte_fallback=True(256 byte pieces). - Corpus: 2,000,000 deduplicated sentences, sentence-level sampled
- 55% Chinese —
HuggingFaceFW/fineweb-2(cmn_Hani) - 45% English —
HuggingFaceFW/fineweb(sample-10BT) - corpus sha256:
4ec6384ada54e71aadc7964f0f0f3aabf0e6d43131bad1766cfcb2ad404448df(dataset revisions are not pinned, so re-training yields a different corpus).
Evaluation (held-out sample sentences)
(Reference: an earlier 8K tokenizer scored zh ≈ 1.0, en ≈ 1.09 chars/token.)
Data statement
Trained on text from the FineWeb / FineWeb-2 datasets (ODC-BY). The tokenizer contains no verbatim documents — only subword statistics.
