CoolFace
Datasetpublic

Ba2han/long_corpus-0209_tokenized

long_corpus-0209_tokenized Exact-deduplicated, tokenized, and shuffled pretraining mix. Processing Tokenizer: /workspace/good_tokenizer (vocab_size=60800) Special tokens: BOS=<bos> (id 2), EOS=<eos> (id 3) bos_eos = True: every example is [BOS] + content + [EOS] min_tokens = 15 (including BOS/EOS) max_tokens = 3072: content truncated to 3070 tokens, then BOS+EOS appended (3070 + eos + bos) Dedup: exact match on stripped UTF-8 text (blake2s-128) Shuffle:… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/long_corpus-0209_tokenized.

sourceHugging Faceotherupdated 23d agoView on Hugging Face
0likes213downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
Ba2han/long_corpus-0209_tokenized · CoolFace