CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01marin-community /fineweb-edu-pretokenized-10K Marin/Levanter Subsampled Pretokenized Dataset Dataset Train Urls: gs://marin-us-central2/raw/fineweb-edu-c2beb4/3c452cb/huggingface.co/datasets/HuggingFaceFW/fineweb-edu/resolve/3c452cb Factsheet Original cache: gs://marin-us-central2/tokenized/fineweb-edu-24698d Tokenizer: stanford-crfm/marin-tokenizer Seed 42 Number of tokens: 10,390 (This readme is automatically generated by Marin.) 0 likes4k downloads1y agoHugging Face02pico-lm /pretokenized-dolma The Pretokenized Dolma Dataset A pre-tokenized, pre-shuffled version of Dolma, the high-quality text corpus from AI2. This dataset is designed to be plug-and-play with the pico-train library. Overview Key Features: Tokenized with allenai/OLMo-7B-0724-hf, a BPE-tokenized with a vocabulary size of 50280 Sequence length: 2049 tokens (2048 + 1 for next-token prediction) Sharded into 10,000 Parquet files (~78MB each) 420B tokens total size (perfect for training a model for… See the full description on the dataset page: https://huggingface.co/datasets/pico-lm/pretokenized-dolma.100M<n<1B4 likes3k downloads1y agoHugging Face03afg1 /rnacentral-pretokenized-sharded10M<n<100M0 likes740 downloads3mo agoHugging Face04leonidas123 /gemma-4-pretokenized-traces1M<n<10M0 likes694 downloads5mo agoHugging Face05jensjepsen /danish-pretokenized-16k Danish Pretrain — pretokenized (byte-level BPE, 16k vocab) — v1 partial Pretokenized version of jensjepsen/danish-pretrain via jensjepsen/danish-tokenizer. Note: This is a partial version — 94 of 96 shards. The 2 missing shards crashed during tokenization (rare pathological content). See jensjepsen/danish-pretokenized-16k-supp for the recovered rows. 93,200,464 docs Schema: input_ids: list<int32>, attention_mask: list<int8> 94 zstd-compressed parquet shards 10M<n<100M0 likes661 downloads2mo agoHugging Face06cskokgibbs /yeast-no-GTL-pretokenized-NT-v2tabular1M<n<10M0 likes414 downloads1y agoHugging Face07AIGym /pretokenized-pretrain-test-128K1M<n<10M0 likes394 downloads1y agoHugging Face08ThomasTheMaker /pretokenized-dolma-5M1M<n<10M0 likes356 downloads1y agoHugging Face09cskokgibbs /mouse-pretokenized-NTtabular100K<n<1M0 likes310 downloads2y agoHugging Face10sigil-ml /PreTokenizedWikiEntext1M<n<10M0 likes306 downloads2y agoHugging Face11cskokgibbs /human-and-mouse-pretokenized-NT-cache0 likes305 downloads2y agoHugging Face12cskokgibbs /human-pretokenized-NTtabular1M<n<10M0 likes304 downloads2y agoHugging Face13ThomasTheMaker /pretokenized-dolma-20M10M<n<100M0 likes300 downloads1y agoHugging Face14cskokgibbs /yeast-gene-sequence-homology-pretokenized-NTtabular1M<n<10M0 likes274 downloads1y agoHugging Face15ZhuofengLi /fineweb-edu-pretokenized-llama3-100b FineWeb-Edu Pretokenized with Llama 3.1 (100B) This repository contains the sample/100BT subset of FineWeb-Edu, pretokenized for Megatron-LM-style training with meta-llama/Meta-Llama-3.1-8B. It is a derived, pretokenized version of the upstream data, not an official Hugging Face FineWeb release. Dataset summary 140 indexed shards 97,270,686 non-empty documents 97,458,793,013 tokens English web text from FineWeb-Edu sample/100BT Source dataset revision:… See the full description on the dataset page: https://huggingface.co/datasets/ZhuofengLi/fineweb-edu-pretokenized-llama3-100b.tabularn<1K0 likes261 downloads2mo agoHugging Face16cskokgibbs /yeast-no-label-pretokenized-NTtabular1M<n<10M0 likes242 downloads1y agoHugging Face17cskokgibbs /human_and_mouse-pretokenized-NTtabular1M<n<10M0 likes239 downloads2y agoHugging Face18marin-community /fineweb-edu-pretokenized-10B Marin/Levanter Subsampled Pretokenized Dataset Dataset Train Urls: gs://marin-us-central2/raw/fineweb-edu-c2beb4/3c452cb/huggingface.co/datasets/HuggingFaceFW/fineweb-edu/resolve/3c452cb Factsheet Original cache: gs://marin-us-central2/tokenized/fineweb-edu-24698d Tokenizer: stanford-crfm/marin-tokenizer Seed 42 Number of tokens: 10,000,000,738 (This readme is automatically generated by Marin.) 0 likes225 downloads1y agoHugging Face19cskokgibbs /BEELINE-HepG2-no-GTL-pretokenized-NTtabular1M<n<10M0 likes224 downloads1y agoHugging Face20cskokgibbs /BEELINE-mDC-no-label-pretokenized-NTtabular1M<n<10M0 likes223 downloads1y agoHugging Face21cskokgibbs /BEELINE-human-v2-pretokenized-NTtabular10M<n<100M0 likes211 downloads1y agoHugging Face22cskokgibbs /BEELINE-mouse-v2-pretokenized-NTtabular1M<n<10M0 likes211 downloads1y agoHugging Face23cskokgibbs /BEELINE-mESC-no-label-pretokenized-NTtabular1M<n<10M0 likes202 downloads1y agoHugging Face24cskokgibbs /BEELINE-mHSC-no-label-pretokenized-NTtabular1M<n<10M0 likes188 downloads1y agoHugging Face25ZhuofengLi /pretraining-pretokenized-smollm3 SmolLM3 Pretokenized Pretraining Sources Datatrove/Nanotron tokenized-byte versions of three public pretraining sources: fineweb-edu-10bt: HuggingFaceFW/fineweb-edu, sample/10BT finemath-4plus: HuggingFaceTB/finemath, finemath-4plus stack-edu-python: HuggingFaceTB/stack-edu, Python, with content retrieved from the public Software Heritage S3 bucket using the upstream dataset-card procedure All subsets use HuggingFaceTB/SmolLM3-3B at revision… See the full description on the dataset page: https://huggingface.co/datasets/ZhuofengLi/pretraining-pretokenized-smollm3.tabularn<1K0 likes161 downloads2mo agoHugging Face26rijuludar /100B-pretokenized-mix-32k --- license: apache-2.0 task_categories: - text-generation language: - en tags: - pretokenized - slm-tokenizer-32k - dclm - fineweb-edu - cosmopedia - math - finemath - python - code - wikipedia - educational - web - stem size_categories: - 100B<n<1T --- SLM Pre-tokenized 100B Mix (32k Vocab) This repository contains a pre-tokenized, multi-domain dataset mixture of approximately 100 Billion tokens across educational, web, encyclopedic, and… See the full description on the dataset page: https://huggingface.co/datasets/rijuludar/100B-pretokenized-mix-32k.0 likes158 downloads3mo agoHugging Face27chankhavu /smolmo-sft-olmocore-pretokenized0 likes148 downloads4mo agoHugging Face28cskokgibbs /BEELINE-hESC-no-GTL-pretokenized-NTtabular1M<n<10M0 likes135 downloads1y agoHugging Face29cskokgibbs /BEELINE-mDC-no-GTL-pretokenized-NTtabular1M<n<10M0 likes133 downloads1y agoHugging Face30ZhuofengLi /fineweb-edu-pretokenized-10b FineWeb-Edu Pretokenized 10B Megatron indexed-dataset (.bin/.idx) versions of HuggingFaceFW/fineweb-edu sample/10BT. Each Hugging Face subset contains a small metadata.parquet index. The actual Megatron files are under <subset>/files/; pass each prefix without the .bin/.idx suffix to Megatron Core or Megatron Bridge. Documents retain upstream shard and row order. Tokenization disables automatic special-token insertion and appends exactly one tokenizer EOS/EOD token to each… See the full description on the dataset page: https://huggingface.co/datasets/ZhuofengLi/fineweb-edu-pretokenized-10b.tabularn<1K0 likes108 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.