CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceFW /finepdfs_50BT-dclm_30BT-fineweb_edu_20BT FinePDFs 50BT + DCLM 30BT + FineWeb-Edu 20BT A ~100 billion token pretraining mixture combining three high-quality English data sources in a 50-30-20 ratio. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Inspired by optimal dataset mixing. Dataset Description Component Source Tokens FinePDFs 100BT FinePDFs ~50B DCLM 100BT DCLM-Baseline 1.0 ~30B FineWeb-Edu 100BT FineWeb-Edu ~20B The schema is reduced to the… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_50BT-dclm_30BT-fineweb_edu_20BT.text10M<n<100M2 likes2.9k downloads7mo agoHugging Face02manu /code-20b Dataset Card for "code_20b2" More Information needed text10M<n<100M4 likes2.8k downloads3y agoHugging Face03HuggingFaceFW /finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT FinePDFs-Edu 50BT + DCLM 30BT + FineWeb-Edu 20BT A ~100 billion token pretraining mixture combining three high-quality English data sources in a 50-30-20 ratio, using the educational subset of FinePDFs. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Inspired by optimal dataset mixing. Dataset Description Component Source Tokens FinePDFs-Edu 100BT FinePDFs-Edu ~50B DCLM 100BT DCLM-Baseline 1.0 ~30B FineWeb-Edu 100BT… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT.text10M<n<100M0 likes2.3k downloads7mo agoHugging Face04Fredithefish /Nemotron-CC-HQ-20B Nemotron-CC-HQ-20B This Dataset consists of approximately 20B tokens of Nemotron-CC-HQ, consisting of randomly sampled slices from crawls in the range CC-MAIN-2013-20-part-00012 to CC-MAIN-2019-04-part-00007. For more information about Nemotron-CC check the Paper by Nvidia Disclaimer: Derived from Nemotron-CC (Common Crawl). No ownership of underlying content is claimed. Data may be subject to third-party rights. Use at your own risk and in compliance with… See the full description on the dataset page: https://huggingface.co/datasets/Fredithefish/Nemotron-CC-HQ-20B.texttext-generation10M<n<100M1 likes2.3k downloads6mo agoHugging Face05manu /code_20b Dataset Card for "code_20b" More Information needed text10M<n<100M1 likes2k downloads3y agoHugging Face06ByteSpanTokenisers /finewebedu-20B FineWebEDU 20B A copy of FineWebEDU-20B used for out tokenizer experiments. The subsets are as follows: bytelevel: the full dataset tokenized using our bytelevel tokenizer bytelevel-subset_1: a 100k-row subset of the bytelevel subset, used to train bytelevel models. bytelevel-subset_2: a 100k-row subset of the bytelevel subset, used to extract llm predictions. bytelevel-llm-data: a copy of bytelevel-subset_2 with lm predictions, used to train bytespan tokenizers… See the full description on the dataset page: https://huggingface.co/datasets/ByteSpanTokenisers/finewebedu-20B.text100M<n<1B0 likes1.7k downloads1y agoHugging Face07HuggingFaceFW /finepdfs_50BT-dclm_30BT-fineweb_edu_20BT-shuffled FinePDFs 50BT + DCLM 30BT + FineWeb-Edu 20BT (Shuffled) A globally shuffled version of HuggingFaceFW/finepdfs_50BT-dclm_30BT-fineweb_edu_20BT. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B token mixture (50B FinePDFs + 30B DCLM + 20B FineWeb-Edu) but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining —… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_50BT-dclm_30BT-fineweb_edu_20BT-shuffled.text10M<n<100M6 likes1.3k downloads7mo agoHugging Face08HuggingFaceFW /finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT-shuffled FinePDFs-Edu 50BT + DCLM 30BT + FineWeb-Edu 20BT (Shuffled) A globally shuffled version of HuggingFaceFW/finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B token mixture (50B FinePDFs-Edu + 30B DCLM + 20B FineWeb-Edu) but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT-shuffled.text10M<n<100M2 likes1.2k downloads7mo agoHugging Face09orionweller /dolma_20bn_wiki_upsampletabular10M<n<100M0 likes997 downloads2y agoHugging Face10alexkstern /fineweb-nanochatbpe-20B fineweb-nanochatbpe-20B FineWeb-Edu (from karpathy/fineweb-edu-100b-shuffle), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536) and packaged as flat uint16 token-id .bin files for fast memmap training. file split tokens train.bin train 20,000,000,000 val.bin val 52,336,096 train and val are disjoint held-out partitions. Each .bin is a raw little-endian uint16 stream (no header); token count = filesize / 2, and train.meta.json / val.meta.jsoncarry the full… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/fineweb-nanochatbpe-20B.tabularn<1K0 likes866 downloads4mo agoHugging Face11andyrdt /gpt-oss-20b-rollouts GPT-OSS-20B Rollouts Generated rollouts from GPT-OSS-20B with parsed Harmony channels (assistant thinking/final). Schema: user_content, system_reasoning_effort, assistant_thinking, assistant_content. Loading example: load_dataset("andyrdt/gpt-oss-20b-rollouts", "HarmBench", split="standard_train"). Notes This repository uses manual configuration to expose both subset (config) and split dropdowns in the viewer. Safety and jailbreak HarmBench: Safety prompts… See the full description on the dataset page: https://huggingface.co/datasets/andyrdt/gpt-oss-20b-rollouts.text1M<n<10M7 likes706 downloads9mo agoHugging Face12jacobmorrison /gpt-oss-20b-combined-outputstext1M<n<10M0 likes688 downloads1y agoHugging Face13orionweller /dolma_20bn_instruct_upsampletabular10M<n<100M0 likes662 downloads2y agoHugging Face14luowenyang /the-stack-v2-20B-sampletext1M<n<10M0 likes535 downloads1y agoHugging Face15project-telos /gpt_oss_20b_doorkey_boundary_activationstabularn<1K0 likes535 downloads3mo agoHugging Face16DanielGallagherIRE /fineweb-edu-20B-E1-Scramble-Hier-flattenedtext1M<n<10M0 likes517 downloads2mo agoHugging Face17orionweller /dolma_20bn_cc_high_qualitytabular10M<n<100M0 likes448 downloads2y agoHugging Face18orionweller /dolma_20bn_prop_stratified_sampletabular10M<n<100M0 likes433 downloads2y agoHugging Face19EleutherAI /SmolLM2-1.7B-stage-4-20Btext10M<n<100M0 likes422 downloads1y agoHugging Face20dahara1 /FineWeb2-HQ-ja-20B元のデータセットFineWeb2-HQ 元のデータセットは多言語で巨大なため、扱いやすい用に日本語データを約200GBだけ抽出したデータセットです wc 結果 1763269 38541549 5370473709 fineweb_jpn_Jpan_chunk_0.jsonl 1784158 37430170 5370514369 fineweb_jpn_Jpan_chunk_1.jsonl 1639554 40065129 5370372344 fineweb_jpn_Jpan_chunk_10.jsonl 1575127 42167166 5370298354 fineweb_jpn_Jpan_chunk_11.jsonl 1686375 39225898 5370402506 fineweb_jpn_Jpan_chunk_12.jsonl 1786948 36456352 5370498572… See the full description on the dataset page: https://huggingface.co/datasets/dahara1/FineWeb2-HQ-ja-20B.text10M<n<100M0 likes399 downloads1y agoHugging Face21moca-embed /dclm_20b DCLM used in MoCa Pre-training 🏠 Homepage | 💻 Code | 🤖 MoCa-Qwen25VL-7B | 🤖 MoCa-Qwen25VL-3B | 📚 Datasets | 📄 Paper Introduction This is a text pre-training dataset used in the modality-aware continual pre-training of MoCa models. It is adapted from DCLM and randomly downsampled to ~20B tokens. The dataset consists of text examples. text is a string containing text while images are left blank intentionally since there is no image available. Citation… See the full description on the dataset page: https://huggingface.co/datasets/moca-embed/dclm_20b.text10M<n<100M0 likes381 downloads1y agoHugging Face22manu /old_code_20b_separatetext1M<n<10M0 likes380 downloads3y agoHugging Face23DanielGallagherIRE /fineweb-edu-20btext10M<n<100M0 likes363 downloads4mo agoHugging Face24pietrolesci /finewebedu-20BThis is a subset of the HuggingFaceFW/fineweb-edu/100BT dataset. I extracted (in order) the initial 20,200,000 rows where, ideally, 20M are meant for training and 200k for validation. Tokenised configs: bpe32000minipile: 21.6B tokens License For the license, refer to the original dataset (HuggingFaceFW/fineweb-edu). tabulartext-generation10M<n<100M1 likes341 downloads2y agoHugging Face25orionweller /dolma_20bn_no_instructtabular10M<n<100M0 likes324 downloads2y agoHugging Face26DanielGallagherIRE /FineWeb-Edu-20B-E2-v3-Obfuscatedtext1M<n<10M0 likes318 downloads1mo agoHugging Face27ericrcwu /week1-general-20b-dolma2-v1 Week-One General 20B Dolma2 This is a deterministic, pretokenized 20-billion-token baseline corpus for controlled language-model architecture and training experiments. It contains nested 100M, 1B, 5B, and 20B views; each larger view is an exact ordered extension of the previous view. It also includes a dataset-only 370m-1.25xc view: the first 9,281,564,672 packed tokens of the verified 20B order, sized for the 1.25xC target of the OLMo-ladder 370M parameter count. The repository… See the full description on the dataset page: https://huggingface.co/datasets/ericrcwu/week1-general-20b-dolma2-v1.texttext-generationn<1K0 likes312 downloads2mo agoHugging Face28DanielGallagherIRE /FineWeb-Edu-20B-E2-OBF-FULLtext1M<n<10M0 likes306 downloads1mo agoHugging Face29PleIAs /gpt-oss20b-samples-dedupA simple deduplicated variant of https://huggingface.co/datasets/jxm/gpt-oss20b-samples Given the predictability of synthetic data we opted for a simple strategy: keeping the unique combinations of first and last ten words. Total count of unique occurrences is available in the column occurrence_count. text100K<n<1M5 likes302 downloads1y agoHugging Face30DanielGallagherIRE /FineWeb-Edu-20B-E2-v3text1M<n<10M0 likes287 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.