CoolFace
Modelpublic

tachiwin/tokenizer_test3

sourceHugging Faceupdated 27d agoView on Hugging Face
0likes
Model Card

Tachiwin multilingual tokenizer

A multilingual ByteLevel-BPE tokenizer trained with Hugging Face Tokenizers.

Corpus weighting

ComponentTarget
Modern + old exotic-language data70%
English10%
Spanish10%
Code10%

The complete available exotic-language corpus is used as the 70% anchor. Its existing modern/old composition is preserved.

Tokenizer

  • Model: BPE, trained via Tokenizer.train_from_iterator over a streaming line generator (constant memory regardless of corpus size)
  • Vocabulary target: 256,000
  • Initial alphabet: complete ByteLevel alphabet
  • ByteLevel GPT-2 regex: disabled
  • Unicode normalizer: none
  • Special tokens: 126
  • Human-language tags: 92

Corpus size

Total materialized corpus: 115,570,108 bytes (0.108 GiB)

Evaluation

The tokenizer was evaluated against the Tachiwin language catalogue. Languages without text samples are skipped. For each language, all available samples are concatenated ONLY within that language for aggregate fertility statistics. Metrics: characters/token, tokens/character, UTF-8 bytes/token, tokens/UTF-8 byte, exact round-trip preservation.

Important training note

The Hugging Face BPE trainer does not expose an internal resumable merge-state checkpoint. The recipe therefore treats the completed tokenizer.json as the training checkpoint:

  • corpus preparation is resumable (both the exotic streaming pass and the capped external-corpus downloads reuse existing shards);
  • recipe/statistics/checksums are stored in recipe/;
  • if tokenizer.json already exists, subsequent runs skip BPE training;
  • an interrupted BPE computation itself must be restarted.

Repository evaluation artifacts

  • evaluation/catalogue.json
  • evaluation/language_fertility.csv
  • evaluation/language_fertility.json
  • evaluation/evaluation_summary.json