kacperwikiel/speakleash-tokenizer-5gb-sample
SpeakLeash tokenizer 42GB quality sample Private tokenizer-training sample built from a stratified local SpeakLeash text snapshot on gb10. Target: 5.0 GiB raw JSONL UTF-8 bytes after document filters and exact text dedup. Format: data/tokenizer_sample_*.jsonl.zst, one JSON object per line with text, source, source_kind. This is intended for tokenizer/BPE training convergence tests.
066
SpeakLeash tokenizer 42GB quality sample
Private tokenizer-training sample built from a stratified local SpeakLeash text snapshot on gb10.
Target: 5.0 GiB raw JSONL UTF-8 bytes after document filters and exact text dedup. Format: data/tokenizer_sample_*.jsonl.zst, one JSON object per line with text, source, source_kind.
This is intended for tokenizer/BPE training convergence tests.
