CoolFace
Datasetpublic

kacperwikiel/speakleash-tokenizer-5gb-sample

SpeakLeash tokenizer 42GB quality sample Private tokenizer-training sample built from a stratified local SpeakLeash text snapshot on gb10. Target: 5.0 GiB raw JSONL UTF-8 bytes after document filters and exact text dedup. Format: data/tokenizer_sample_*.jsonl.zst, one JSON object per line with text, source, source_kind. This is intended for tokenizer/BPE training convergence tests.

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
0likes66downloads
Dataset Card

SpeakLeash tokenizer 42GB quality sample

Private tokenizer-training sample built from a stratified local SpeakLeash text snapshot on gb10.

Target: 5.0 GiB raw JSONL UTF-8 bytes after document filters and exact text dedup. Format: data/tokenizer_sample_*.jsonl.zst, one JSON object per line with text, source, source_kind.

This is intended for tokenizer/BPE training convergence tests.