CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01pico-lm /pretokenized-dolma The Pretokenized Dolma Dataset A pre-tokenized, pre-shuffled version of Dolma, the high-quality text corpus from AI2. This dataset is designed to be plug-and-play with the pico-train library. Overview Key Features: Tokenized with allenai/OLMo-7B-0724-hf, a BPE-tokenized with a vocabulary size of 50280 Sequence length: 2049 tokens (2048 + 1 for next-token prediction) Sharded into 10,000 Parquet files (~78MB each) 420B tokens total size (perfect for training a model for… See the full description on the dataset page: https://huggingface.co/datasets/pico-lm/pretokenized-dolma.100M<n<1B4 likes3k downloads1y agoHugging Face02afg1 /rnacentral-pretokenized-sharded10M<n<100M0 likes740 downloads3mo agoHugging Face03leonidas123 /gemma-4-pretokenized-traces1M<n<10M0 likes694 downloads5mo agoHugging Face04jensjepsen /danish-pretokenized-16k Danish Pretrain — pretokenized (byte-level BPE, 16k vocab) — v1 partial Pretokenized version of jensjepsen/danish-pretrain via jensjepsen/danish-tokenizer. Note: This is a partial version — 94 of 96 shards. The 2 missing shards crashed during tokenization (rare pathological content). See jensjepsen/danish-pretokenized-16k-supp for the recovered rows. 93,200,464 docs Schema: input_ids: list<int32>, attention_mask: list<int8> 94 zstd-compressed parquet shards 10M<n<100M0 likes661 downloads2mo agoHugging Face05cskokgibbs /yeast-no-GTL-pretokenized-NT-v2tabular1M<n<10M0 likes414 downloads1y agoHugging Face06AIGym /pretokenized-pretrain-test-128K1M<n<10M0 likes394 downloads1y agoHugging Face07ThomasTheMaker /pretokenized-dolma-5M1M<n<10M0 likes356 downloads1y agoHugging Face08cskokgibbs /mouse-pretokenized-NTtabular100K<n<1M0 likes310 downloads2y agoHugging Face09cskokgibbs /human-pretokenized-NTtabular1M<n<10M0 likes304 downloads2y agoHugging Face10ThomasTheMaker /pretokenized-dolma-20M10M<n<100M0 likes300 downloads1y agoHugging Face11cskokgibbs /yeast-gene-sequence-homology-pretokenized-NTtabular1M<n<10M0 likes274 downloads1y agoHugging Face12cskokgibbs /yeast-no-label-pretokenized-NTtabular1M<n<10M0 likes242 downloads1y agoHugging Face13cskokgibbs /human_and_mouse-pretokenized-NTtabular1M<n<10M0 likes239 downloads2y agoHugging Face14cskokgibbs /BEELINE-HepG2-no-GTL-pretokenized-NTtabular1M<n<10M0 likes224 downloads1y agoHugging Face15cskokgibbs /BEELINE-mDC-no-label-pretokenized-NTtabular1M<n<10M0 likes223 downloads1y agoHugging Face16cskokgibbs /BEELINE-human-v2-pretokenized-NTtabular10M<n<100M0 likes211 downloads1y agoHugging Face17cskokgibbs /BEELINE-mouse-v2-pretokenized-NTtabular1M<n<10M0 likes211 downloads1y agoHugging Face18cskokgibbs /BEELINE-mESC-no-label-pretokenized-NTtabular1M<n<10M0 likes202 downloads1y agoHugging Face19cskokgibbs /BEELINE-mHSC-no-label-pretokenized-NTtabular1M<n<10M0 likes188 downloads1y agoHugging Face20ZhuofengLi /pretraining-pretokenized-smollm3 SmolLM3 Pretokenized Pretraining Sources Datatrove/Nanotron tokenized-byte versions of three public pretraining sources: fineweb-edu-10bt: HuggingFaceFW/fineweb-edu, sample/10BT finemath-4plus: HuggingFaceTB/finemath, finemath-4plus stack-edu-python: HuggingFaceTB/stack-edu, Python, with content retrieved from the public Software Heritage S3 bucket using the upstream dataset-card procedure All subsets use HuggingFaceTB/SmolLM3-3B at revision… See the full description on the dataset page: https://huggingface.co/datasets/ZhuofengLi/pretraining-pretokenized-smollm3.tabularn<1K0 likes161 downloads2mo agoHugging Face21cskokgibbs /BEELINE-hESC-no-GTL-pretokenized-NTtabular1M<n<10M0 likes135 downloads1y agoHugging Face22cskokgibbs /BEELINE-mDC-no-GTL-pretokenized-NTtabular1M<n<10M0 likes133 downloads1y agoHugging Face23ZhuofengLi /fineweb-edu-pretokenized-10b FineWeb-Edu Pretokenized 10B Megatron indexed-dataset (.bin/.idx) versions of HuggingFaceFW/fineweb-edu sample/10BT. Each Hugging Face subset contains a small metadata.parquet index. The actual Megatron files are under <subset>/files/; pass each prefix without the .bin/.idx suffix to Megatron Core or Megatron Bridge. Documents retain upstream shard and row order. Tokenization disables automatic special-token insertion and appends exactly one tokenizer EOS/EOD token to each… See the full description on the dataset page: https://huggingface.co/datasets/ZhuofengLi/fineweb-edu-pretokenized-10b.tabularn<1K0 likes108 downloads2mo agoHugging Face24nomic-ai /nomic-bert-pretokenized-2048-wiki-20231M<n<10M0 likes91 downloads2y agoHugging Face25EleutherAI /tinystories-pretokenized-pythia Dataset Card for "tinystories-pretokenized-pythia" More Information needed 100K<n<1M0 likes79 downloads2y agoHugging Face26cskokgibbs /yeast-tf-sequence-homology-pretokenized-NTtabular1M<n<10M0 likes77 downloads1y agoHugging Face27cskokgibbs /BEELINE-mESC-no-GTL-pretokenized-NTtabular100K<n<1M0 likes71 downloads1y agoHugging Face28cskokgibbs /BEELINE-HepG2-no-label-pretokenized-NTtabular100K<n<1M0 likes67 downloads1y agoHugging Face29cskokgibbs /BEELINE-hESC-no-label-pretokenized-NTtabular100K<n<1M0 likes67 downloads1y agoHugging Face30AWfaw /ai-hdlcoder-pretokenized-dataset-train Dataset Card for "ai-hdlcoder-pretokenized-dataset-train" More Information needed 10K<n<100K0 likes62 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.