Ba2han/long_corpus-0209_tokenized
long_corpus-0209_tokenized Exact-deduplicated, tokenized, and shuffled pretraining mix. Processing Tokenizer: /workspace/good_tokenizer (vocab_size=60800) Special tokens: BOS=<bos> (id 2), EOS=<eos> (id 3) bos_eos = True: every example is [BOS] + content + [EOS] min_tokens = 15 (including BOS/EOS) max_tokens = 3072: content truncated to 3070 tokens, then BOS+EOS appended (3070 + eos + bos) Dedup: exact match on stripped UTF-8 text (blake2s-128) Shuffle:… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/long_corpus-0209_tokenized.
longcorpus-0209tokenized
Exact-deduplicated, tokenized, and shuffled pretraining mix.
Processing
- Tokenizer:
/workspace/good_tokenizer(vocab_size=60800) - Special tokens: BOS=
<bos>(id 2), EOS=<eos>(id 3) bos_eos = True: every example is[BOS] + content + [EOS]min_tokens = 15(including BOS/EOS)max_tokens = 3072: content truncated to 3070 tokens, then BOS+EOS appended (3070 + eos + bos)- Dedup: exact match on stripped UTF-8 text (
blake2s-128) - Shuffle: hash-partition into 128 buckets, permute buckets and rows with seed
42 - Columns:
input_ids(list<int32>),token_count(int32) - Parquet: zstd, ~100,000 rows/shard
Sources were streamed interleaved. Documents that collide on the exact-text hash keep the first copy seen in interleave order.
Token statistics
Per-source counts
Sources
Ba2han/long_corpus-0209— primary long mix (TR/AZ/CRH/TUK/EN)Ba2han/temiz-OSCAR-long-all— filtered long OSCARBa2han/finepdfs-long— filtered long FinePDFs (tur_Latn)leukas/climbmix-1b-shuffle— 1B-word ClimbMix sample
Source licenses still apply. long_corpus-0209 is mixed/other (includes CC BY-NC and non-commercial derived rows).
