Ba2han/long_corpus-0209_tokenized
long_corpus-0209_tokenized Exact-deduplicated, tokenized, and shuffled pretraining mix. Processing Tokenizer: /workspace/good_tokenizer (vocab_size=60800) Special tokens: BOS=<bos> (id 2), EOS=<eos> (id 3) bos_eos = True: every example is [BOS] + content + [EOS] min_tokens = 15 (including BOS/EOS) max_tokens = 3072: content truncated to 3070 tokens, then BOS+EOS appended (3070 + eos + bos) Dedup: exact match on stripped UTF-8 text (blake2s-128) Shuffle:… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/long_corpus-0209_tokenized.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face