CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceFW /fineweb_edu_100BT-shuffled FineWeb-Edu 100BT (Shuffled) A globally shuffled version of HuggingFaceFW/fineweb_edu_100BT. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B tokens as fineweb_edu_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining. How It Was Created The unshuffled dataset was loaded into memory, shuffled… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb_edu_100BT-shuffled.tabular100M<n<1B6 likes4.3k downloads7mo agoHugging Face02DanielGallagherIRE /FineWeb-Edu-10B-Shuffledtext1M<n<10M0 likes2k downloads3mo agoHugging Face03Muesli1 /dclm-baseline-1.0-llama3-tokenized-shuffled !! Note: this dataset is currently being uploaded and processed. The .bin files are intermediate files to allow shuffling. !! DCLM-Baseline Pretokenized (LLaMA 3.1, 8192 context) This dataset is a pretokenized and globally shuffled version of DCLM-Baseline (mlfoundations/dclm-baseline-1.0), prepared for large-scale language model pretraining. It is intended to be used as a direct drop-in pretraining corpus for LLaMA 3.1 style training pipelines. The original DCLM-Baseline… See the full description on the dataset page: https://huggingface.co/datasets/Muesli1/dclm-baseline-1.0-llama3-tokenized-shuffled.text0 likes1.8k downloads5mo agoHugging Face04HuggingFaceFW /finepdfs_50BT-dclm_30BT-fineweb_edu_20BT-shuffled FinePDFs 50BT + DCLM 30BT + FineWeb-Edu 20BT (Shuffled) A globally shuffled version of HuggingFaceFW/finepdfs_50BT-dclm_30BT-fineweb_edu_20BT. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B token mixture (50B FinePDFs + 30B DCLM + 20B FineWeb-Edu) but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining —… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_50BT-dclm_30BT-fineweb_edu_20BT-shuffled.text10M<n<100M6 likes1.4k downloads7mo agoHugging Face05HuggingFaceFW /dclm_100BT-shuffled DCLM 100BT (Shuffled) A globally shuffled version of HuggingFaceFW/dclm_100BT. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B tokens as dclm_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining. How It Was Created The unshuffled dataset was loaded into memory, shuffled with… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/dclm_100BT-shuffled.tabular10M<n<100M3 likes1.4k downloads7mo agoHugging Face06HuggingFaceFW /finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT-shuffled FinePDFs-Edu 50BT + DCLM 30BT + FineWeb-Edu 20BT (Shuffled) A globally shuffled version of HuggingFaceFW/finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B token mixture (50B FinePDFs-Edu + 30B DCLM + 20B FineWeb-Edu) but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT-shuffled.text10M<n<100M2 likes1.3k downloads7mo agoHugging Face07Muesli1 /dclm-baseline-1.0-llama3-tokenized-shuffled-524K !! Note: this dataset is currently being uploaded and processed. The .bin files are intermediate files to allow shuffling. !! DCLM-Baseline Pretokenized (LLaMA 3.1, 524288 context) This dataset is a pretokenized and globally shuffled version of DCLM-Baseline (mlfoundations/dclm-baseline-1.0), prepared for large-scale language model pretraining. It is intended to be used as a direct drop-in pretraining corpus for LLaMA 3.1 style training pipelines. The original… See the full description on the dataset page: https://huggingface.co/datasets/Muesli1/dclm-baseline-1.0-llama3-tokenized-shuffled-524K.text0 likes1k downloads5mo agoHugging Face08RedMod /MidTool-Mix-shuffledshuffled all sets MidTool-Mix A 20.3B-token mid-training corpus for agentic tool use. It pairs filtered web, PDF, and code sources with synthesized agent supervision, and is designed to teach models to recognize tool affordances, ground arguments from context, compose tool-call workflows, and recover from incomplete information — before any post-training. Mid-training Qwen3-4B-Base / Qwen3-8B-Base on MidTool-Mix improves downstream tool use under both SFT and RL on BFCLv3… See the full description on the dataset page: https://huggingface.co/datasets/RedMod/MidTool-Mix-shuffled.texttext-generation10M<n<100M0 likes1k downloads29d agoHugging Face09dlwh /MultiLegalPile_Wikipedia_Shuffledtext100K<n<1M0 likes990 downloads4y agoHugging Face10manu /tok-corpus-shuffled Dataset Card for "tok-corpus-shuffled" This is the dataset used to fit custom tokenizers. Goal is to have a ytokenizer that is good for French, Englsih and Code. The dataset uploaded is shuffled to facilitate subsampling it for tokenizer training. French Dataset({ features: ['id', 'text', 'dataset_id'], num_rows: 16881941 }) Code Dataset({ features: ['id', 'text', 'dataset_id'], num_rows: 6338566 }) English Dataset({ features: ['text', 'id', 'dataset_id']… See the full description on the dataset page: https://huggingface.co/datasets/manu/tok-corpus-shuffled.text10M<n<100M1 likes897 downloads3y agoHugging Face11HuggingFaceFW /finepdfs_edu_100BT-shuffled FinePDFs-Edu 100BT (Shuffled) A globally shuffled version of HuggingFaceFW/finepdfs_edu_100BT. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B tokens as finepdfs_edu_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining. How It Was Created The unshuffled dataset was loaded into… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_edu_100BT-shuffled.text10M<n<100M0 likes879 downloads7mo agoHugging Face12HuggingFaceFW /fineweb_100BT-shuffled FineWeb 100BT (Shuffled) A globally shuffled version of HuggingFaceFW/fineweb_100BT. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B tokens as fineweb_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining. How It Was Created The unshuffled dataset was loaded into memory, shuffled with… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb_100BT-shuffled.tabular100M<n<1B0 likes772 downloads7mo agoHugging Face13alea-institute /kl3m-data-sample-004-shuffled KL3M Data Sample 004 (Shuffled) This dataset contains a shuffled sample of 10 million examples from the KL3M Data Project, an initiative by the ALEA Institute providing copyright-clean training resources for large language models across legal, regulatory, and government domains. The KL3M Data Project encompasses approximately 28 TB of compressed documents from authoritative sources including court opinions, government regulatory materials, corporate filings, intellectual property… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-sample-004-shuffled.text10M<n<100M0 likes666 downloads11mo agoHugging Face14redmoddata /nemo23_shuffledtext10M<n<100M0 likes627 downloads7mo agoHugging Face15HuggingFaceFW /finepdfs_100BT-shuffled FinePDFs 100BT (Shuffled) A globally shuffled version of HuggingFaceFW/finepdfs_100BT. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B tokens as finepdfs_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining. How It Was Created The unshuffled dataset was loaded into memory… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_100BT-shuffled.text10M<n<100M0 likes617 downloads7mo agoHugging Face16medarc /TCGA-12K-parquet-shuffled TCGA-12K Parquet (Shuffled) Attribution This dataset contains 224 x 224 JPEG patches from whole-slide images originally downloaded from The Cancer Genome Atlas (TCGA) that are available in the NCI Genomic Data Commons (GDC) Open Access tier. We mirror and repackage a commonly used ~12k WSI subset in parquet format for ease of training. We exclude patches that did not pass HSV thresholding, following the procedure in Kaiko.AI's Midnight paper. Patches were randomly sampled… See the full description on the dataset page: https://huggingface.co/datasets/medarc/TCGA-12K-parquet-shuffled.tabular10M<n<100M0 likes595 downloads10mo agoHugging Face17JonathanMiddleton /fineweb-edu-dedup-shuffled FineWeb-Edu-Dedup (Globally Shuffled) A uniformly shuffled version of the FineWeb-Edu-Dedup subset from SmolLM-Corpus by HuggingFace. Source Data This dataset is derived from HuggingFaceTB/smollm-corpus, specifically the fineweb-edu-dedup subset. That subset is itself derived from FineWeb-Edu, a filtered and deduplicated extract of Common Crawl selected for educational content quality. Property Value Source dataset HuggingFaceTB/smollm-corpus Source subset… See the full description on the dataset page: https://huggingface.co/datasets/JonathanMiddleton/fineweb-edu-dedup-shuffled.texttext-generation100M<n<1B0 likes593 downloads7mo agoHugging Face18nthngdy /mmlu_shuffledtext10K<n<100K0 likes572 downloads1y agoHugging Face19open-r1 /verifiable-coding-problems-python_decontaminated-tested-shuffledtext10K<n<100K2 likes474 downloads1y agoHugging Face20WendyHoang /news-ka-shuffled-ELECTRAtext1M<n<10M0 likes431 downloads2y agoHugging Face21jonathanli /idl-wds-pdfs-images-shuffledtext10M<n<100M0 likes423 downloads10mo agoHugging Face22pingzhili /snli-ve-shuffledimage100K<n<1M0 likes418 downloads3y agoHugging Face23Haesteining /ShuffledDataset2tabular1M<n<10M0 likes410 downloads2y agoHugging Face24sradc /chunked-shuffled-wikipedia20220301en-bookcorpusopen Dataset Card for "wikipedia20220301en-bookcorpusopen-chunked-shuffled" num_examples: 33.5 million download_size: 15.3 GB dataset_size: 26.1 GB This dataset combines wikipedia20220301.en and bookcorpusopen, and splits the data into smaller chunks, of size ~820 chars (such that each item will be at least ~128 tokens for the average tokenizer). The order of the items in this dataset has been shuffled, meaning you don't have to use dataset.shuffle, which is slower to iterate over.… See the full description on the dataset page: https://huggingface.co/datasets/sradc/chunked-shuffled-wikipedia20220301en-bookcorpusopen.text10M<n<100M4 likes404 downloads3y agoHugging Face25hypnopump /smiles_transformer_shuffledtabular100M<n<1B0 likes404 downloads1y agoHugging Face26SandyResearch /fineweb-edu-shuffled FineWeb-EDU Shuffled Pre-shuffled versions of HuggingFaceFW/fineweb-edu. Configs Config Shards ~Rows Description sample-100BT ~1800 ~96M 100B token sample, shuffled sample-350BT ~1340 ~335M 350B token sample, shuffled, deduplicated against val val ~18 ~4.5M Validation set (held out from 100BT) Usage from datasets import load_dataset # Load 100B token training set ds_100b = load_dataset("SandyResearch/fineweb-edu-shuffled"… See the full description on the dataset page: https://huggingface.co/datasets/SandyResearch/fineweb-edu-shuffled.texttext-generation100M<n<1B0 likes402 downloads6mo agoHugging Face27aklein4 /fineweb-edu-sample-10BT-shuffled 📚 FineWeb-Edu (Shuffled) The samples in HuggingFaceFW/fineweb-edu don't appear to be fully shuffled, leading to oscillating loss curves. This dataset contains a shuffled version of the sample-10BT sample from HuggingFaceFW/fineweb-edu. Shuffling was performed using the following script: import datasets data = datasets.load_dataset( "HuggingFaceFW/fineweb-edu", "sample-10BT", split="train", streaming=False, ) data_shuffled = data.shuffle(seed=42)… See the full description on the dataset page: https://huggingface.co/datasets/aklein4/fineweb-edu-sample-10BT-shuffled.tabulartext-generation1M<n<10M1 likes396 downloads1y agoHugging Face28chrisagrams /massive_kb_v1_shuffled MassIVE-KB v1 — globally-shuffled parquet Tandem mass spectra from the MassIVE-KB v1 peptide spectral library, repackaged as globally-shuffled, zstd-compressed Parquet shards for streaming model training. Why this repackaging The source MassIVE-KB MGF is ordered by source raw file, with the same peptide repeated in consecutive entries. Feeding that order to training leaves long runs of correlated spectra inside any locally-shuffled read window, which produces loss… See the full description on the dataset page: https://huggingface.co/datasets/chrisagrams/massive_kb_v1_shuffled.textother10M<n<100M0 likes383 downloads3mo agoHugging Face29TheFinAI /00_filtered_shuffledtabular10M<n<100M0 likes348 downloads4mo agoHugging Face30WendyHoang /news-ka-shuffled-DISTILBERTtext1M<n<10M0 likes341 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.