CoolFace
9 results

pre-tokenized

marin-community /fineweb-edu-pretokenized-10K Marin/Levanter Subsampled Pretokenized Dataset Dataset Train Urls: gs://marin-us-central2/raw/fineweb-edu-c2beb4/3c452cb/huggingface.co/datasets/HuggingFaceFW/fineweb-edu/resolve/3c452cb Factsheet Original cache: gs://marin-us-central2/tokenized/fineweb-edu-24698d Tokenizer: stanford-crfm/marin-tokenizer Seed 42 Number of tokens: 10,390 (This readme is automatically generated by Marin.) 0 likes4k downloads1y agoHugging Facepico-lm /pretokenized-dolma The Pretokenized Dolma Dataset A pre-tokenized, pre-shuffled version of Dolma, the high-quality text corpus from AI2. This dataset is designed to be plug-and-play with the pico-train library. Overview Key Features: Tokenized with allenai/OLMo-7B-0724-hf, a BPE-tokenized with a vocabulary size of 50280 Sequence length: 2049 tokens (2048 + 1 for next-token prediction) Sharded into 10,000 Parquet files (~78MB each) 420B tokens total size (perfect for training a model for… See the full description on the dataset page: https://huggingface.co/datasets/pico-lm/pretokenized-dolma.100M<n<1B4 likes3k downloads1y agoHugging Faceafg1 /rnacentral-pretokenized-sharded10M<n<100M0 likes740 downloads3mo agoHugging Faceleonidas123 /gemma-4-pretokenized-traces1M<n<10M0 likes694 downloads5mo agoHugging Facejensjepsen /danish-pretokenized-16k Danish Pretrain — pretokenized (byte-level BPE, 16k vocab) — v1 partial Pretokenized version of jensjepsen/danish-pretrain via jensjepsen/danish-tokenizer. Note: This is a partial version — 94 of 96 shards. The 2 missing shards crashed during tokenization (rare pathological content). See jensjepsen/danish-pretokenized-16k-supp for the recovered rows. 93,200,464 docs Schema: input_ids: list<int32>, attention_mask: list<int8> 94 zstd-compressed parquet shards 10M<n<100M0 likes661 downloads2mo agoHugging Facecskokgibbs /yeast-no-GTL-pretokenized-NT-v2tabular1M<n<10M0 likes414 downloads1y agoHugging Face