CoolFace
Modelpublic

carry121/k3-32k-tokenizer

sourceHugging Faceodc-byupdated 1mo agoView on Hugging Face
0likes
Model Card

k3-32k tokenizer

Bilingual (Chinese / English) SentencePiece Unigram tokenizer with a 32,000-piece vocabulary, trained for the K3/KQ model family.

Usage

python
from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("carry121/k3-32k-tokenizer")
ids = tok.encode("深度学习是机器学习的一个分支。")
# BOS (<s>, id=2) is prepended automatically (add_bos_token=True).

Both fast (tokenizers backend) and slow (SentencePiece) loading work; encodings are identical.

Special tokens

idpiecerole
0<unk>unknown
1<pad>padding
2<s>BOS (auto-prepended)
3</s>EOS
—<zh2en>, <en2zh>, <source>, <target>user-defined, atomic (reserved for translation / seq2seq-style tasks)

IDs 0–3 are fixed and stable; the four user-defined symbols always encode to a single id.

Training

  • —Algorithm: SentencePiece Unigram, NFKC normalization, character_coverage=0.9995, byte_fallback=True (256 byte pieces).
  • —Corpus: 2,000,000 deduplicated sentences, sentence-level sampled
  • —55% Chinese — HuggingFaceFW/fineweb-2 (cmn_Hani)
  • —45% English — HuggingFaceFW/fineweb (sample-10BT)
  • —corpus sha256: 4ec6384ada54e71aadc7964f0f0f3aabf0e6d43131bad1766cfcb2ad404448df (dataset revisions are not pinned, so re-training yields a different corpus).

Evaluation (held-out sample sentences)

languagecompressionunk rateroundtrip
Chinese1.70 chars/token0✓
English5.02 chars/token0✓

(Reference: an earlier 8K tokenizer scored zh ≈ 1.0, en ≈ 1.09 chars/token.)

Data statement

Trained on text from the FineWeb / FineWeb-2 datasets (ODC-BY). The tokenizer contains no verbatim documents — only subword statistics.