CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceFW /fineweb_edu_100BT-shuffled FineWeb-Edu 100BT (Shuffled) A globally shuffled version of HuggingFaceFW/fineweb_edu_100BT. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B tokens as fineweb_edu_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining. How It Was Created The unshuffled dataset was loaded into memory, shuffled… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb_edu_100BT-shuffled.tabular100M<n<1B6 likes4.2k downloads7mo agoHugging Face02HuggingFaceFW /finepdfs_50BT-dclm_30BT-fineweb_edu_20BT-shuffled FinePDFs 50BT + DCLM 30BT + FineWeb-Edu 20BT (Shuffled) A globally shuffled version of HuggingFaceFW/finepdfs_50BT-dclm_30BT-fineweb_edu_20BT. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B token mixture (50B FinePDFs + 30B DCLM + 20B FineWeb-Edu) but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining —… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_50BT-dclm_30BT-fineweb_edu_20BT-shuffled.text10M<n<100M6 likes1.4k downloads7mo agoHugging Face03HuggingFaceFW /dclm_100BT-shuffled DCLM 100BT (Shuffled) A globally shuffled version of HuggingFaceFW/dclm_100BT. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B tokens as dclm_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining. How It Was Created The unshuffled dataset was loaded into memory, shuffled with… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/dclm_100BT-shuffled.tabular10M<n<100M3 likes1.4k downloads7mo agoHugging Face04HuggingFaceFW /finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT-shuffled FinePDFs-Edu 50BT + DCLM 30BT + FineWeb-Edu 20BT (Shuffled) A globally shuffled version of HuggingFaceFW/finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B token mixture (50B FinePDFs-Edu + 30B DCLM + 20B FineWeb-Edu) but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT-shuffled.text10M<n<100M2 likes1.3k downloads7mo agoHugging Face05manu /tok-corpus-shuffled Dataset Card for "tok-corpus-shuffled" This is the dataset used to fit custom tokenizers. Goal is to have a ytokenizer that is good for French, Englsih and Code. The dataset uploaded is shuffled to facilitate subsampling it for tokenizer training. French Dataset({ features: ['id', 'text', 'dataset_id'], num_rows: 16881941 }) Code Dataset({ features: ['id', 'text', 'dataset_id'], num_rows: 6338566 }) English Dataset({ features: ['text', 'id', 'dataset_id']… See the full description on the dataset page: https://huggingface.co/datasets/manu/tok-corpus-shuffled.text10M<n<100M1 likes899 downloads3y agoHugging Face06HuggingFaceFW /finepdfs_edu_100BT-shuffled FinePDFs-Edu 100BT (Shuffled) A globally shuffled version of HuggingFaceFW/finepdfs_edu_100BT. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B tokens as finepdfs_edu_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining. How It Was Created The unshuffled dataset was loaded into… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_edu_100BT-shuffled.text10M<n<100M0 likes864 downloads7mo agoHugging Face07HuggingFaceFW /fineweb_100BT-shuffled FineWeb 100BT (Shuffled) A globally shuffled version of HuggingFaceFW/fineweb_100BT. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B tokens as fineweb_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining. How It Was Created The unshuffled dataset was loaded into memory, shuffled with… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb_100BT-shuffled.tabular100M<n<1B0 likes762 downloads7mo agoHugging Face08alea-institute /kl3m-data-sample-004-shuffled KL3M Data Sample 004 (Shuffled) This dataset contains a shuffled sample of 10 million examples from the KL3M Data Project, an initiative by the ALEA Institute providing copyright-clean training resources for large language models across legal, regulatory, and government domains. The KL3M Data Project encompasses approximately 28 TB of compressed documents from authoritative sources including court opinions, government regulatory materials, corporate filings, intellectual property… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-sample-004-shuffled.text10M<n<100M0 likes644 downloads11mo agoHugging Face09HuggingFaceFW /finepdfs_100BT-shuffled FinePDFs 100BT (Shuffled) A globally shuffled version of HuggingFaceFW/finepdfs_100BT. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B tokens as finepdfs_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining. How It Was Created The unshuffled dataset was loaded into memory… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_100BT-shuffled.text10M<n<100M0 likes617 downloads7mo agoHugging Face10medarc /TCGA-12K-parquet-shuffled TCGA-12K Parquet (Shuffled) Attribution This dataset contains 224 x 224 JPEG patches from whole-slide images originally downloaded from The Cancer Genome Atlas (TCGA) that are available in the NCI Genomic Data Commons (GDC) Open Access tier. We mirror and repackage a commonly used ~12k WSI subset in parquet format for ease of training. We exclude patches that did not pass HSV thresholding, following the procedure in Kaiko.AI's Midnight paper. Patches were randomly sampled… See the full description on the dataset page: https://huggingface.co/datasets/medarc/TCGA-12K-parquet-shuffled.tabular10M<n<100M0 likes595 downloads10mo agoHugging Face11JonathanMiddleton /fineweb-edu-dedup-shuffled FineWeb-Edu-Dedup (Globally Shuffled) A uniformly shuffled version of the FineWeb-Edu-Dedup subset from SmolLM-Corpus by HuggingFace. Source Data This dataset is derived from HuggingFaceTB/smollm-corpus, specifically the fineweb-edu-dedup subset. That subset is itself derived from FineWeb-Edu, a filtered and deduplicated extract of Common Crawl selected for educational content quality. Property Value Source dataset HuggingFaceTB/smollm-corpus Source subset… See the full description on the dataset page: https://huggingface.co/datasets/JonathanMiddleton/fineweb-edu-dedup-shuffled.texttext-generation100M<n<1B0 likes594 downloads7mo agoHugging Face12nthngdy /mmlu_shuffledtext10K<n<100K0 likes572 downloads1y agoHugging Face13open-r1 /verifiable-coding-problems-python_decontaminated-tested-shuffledtext10K<n<100K2 likes483 downloads2y agoHugging Face14WendyHoang /news-ka-shuffled-ELECTRAtext1M<n<10M0 likes433 downloads2y agoHugging Face15chrisagrams /massive_kb_v1_shuffled MassIVE-KB v1 — globally-shuffled parquet Tandem mass spectra from the MassIVE-KB v1 peptide spectral library, repackaged as globally-shuffled, zstd-compressed Parquet shards for streaming model training. Why this repackaging The source MassIVE-KB MGF is ordered by source raw file, with the same peptide repeated in consecutive entries. Feeding that order to training leaves long runs of correlated spectra inside any locally-shuffled read window, which produces loss… See the full description on the dataset page: https://huggingface.co/datasets/chrisagrams/massive_kb_v1_shuffled.textother10M<n<100M0 likes431 downloads3mo agoHugging Face16pingzhili /snli-ve-shuffledimage100K<n<1M0 likes421 downloads3y agoHugging Face17Haesteining /ShuffledDataset2tabular1M<n<10M0 likes410 downloads2y agoHugging Face18hypnopump /smiles_transformer_shuffledtabular100M<n<1B0 likes410 downloads1y agoHugging Face19jonathanli /idl-wds-pdfs-images-shuffledtext10M<n<100M0 likes407 downloads10mo agoHugging Face20SandyResearch /fineweb-edu-shuffled FineWeb-EDU Shuffled Pre-shuffled versions of HuggingFaceFW/fineweb-edu. Configs Config Shards ~Rows Description sample-100BT ~1800 ~96M 100B token sample, shuffled sample-350BT ~1340 ~335M 350B token sample, shuffled, deduplicated against val val ~18 ~4.5M Validation set (held out from 100BT) Usage from datasets import load_dataset # Load 100B token training set ds_100b = load_dataset("SandyResearch/fineweb-edu-shuffled"… See the full description on the dataset page: https://huggingface.co/datasets/SandyResearch/fineweb-edu-shuffled.texttext-generation100M<n<1B0 likes399 downloads6mo agoHugging Face21aklein4 /fineweb-edu-sample-10BT-shuffled 📚 FineWeb-Edu (Shuffled) The samples in HuggingFaceFW/fineweb-edu don't appear to be fully shuffled, leading to oscillating loss curves. This dataset contains a shuffled version of the sample-10BT sample from HuggingFaceFW/fineweb-edu. Shuffling was performed using the following script: import datasets data = datasets.load_dataset( "HuggingFaceFW/fineweb-edu", "sample-10BT", split="train", streaming=False, ) data_shuffled = data.shuffle(seed=42)… See the full description on the dataset page: https://huggingface.co/datasets/aklein4/fineweb-edu-sample-10BT-shuffled.tabulartext-generation1M<n<10M1 likes396 downloads1y agoHugging Face22sradc /chunked-shuffled-wikipedia20220301en-bookcorpusopen Dataset Card for "wikipedia20220301en-bookcorpusopen-chunked-shuffled" num_examples: 33.5 million download_size: 15.3 GB dataset_size: 26.1 GB This dataset combines wikipedia20220301.en and bookcorpusopen, and splits the data into smaller chunks, of size ~820 chars (such that each item will be at least ~128 tokens for the average tokenizer). The order of the items in this dataset has been shuffled, meaning you don't have to use dataset.shuffle, which is slower to iterate over.… See the full description on the dataset page: https://huggingface.co/datasets/sradc/chunked-shuffled-wikipedia20220301en-bookcorpusopen.text10M<n<100M4 likes395 downloads3y agoHugging Face23TheFinAI /00_filtered_shuffledtabular10M<n<100M0 likes348 downloads4mo agoHugging Face24WendyHoang /news-ka-shuffled-DISTILBERTtext1M<n<10M0 likes341 downloads2y agoHugging Face25another-phytophile /153-angiosperm-species-32k-sequences-shuffledtabular1M<n<10M0 likes317 downloads5mo agoHugging Face26jack-stanley /fineweb-edu-dedup-10b-30gram-shuffledtext1M<n<10M0 likes312 downloads1y agoHugging Face27ching-goodfire /MAPS-ClinVar-VKS-Embeddings-L80-shuffled MAPS ClinVar/VKS ESM-C layer-80 difference fields, shuffled layout The same 200,913 rows as ching-goodfire/MAPS-ClinVar-VKS-Embeddings-L80, rewritten in one random order so that a prefix is a sample. The source repo is sorted by protein accession. That makes the first N rows a block of related proteins rather than a sample of the pool, so any subset that is actually representative requires reading all 154.01 GB and selecting rows afterwards. Here the rows are stored shuffled, so… See the full description on the dataset page: https://huggingface.co/datasets/ching-goodfire/MAPS-ClinVar-VKS-Embeddings-L80-shuffled.tabular100K<n<1M0 likes303 downloads2mo agoHugging Face28konwoo /dclm-16m-shuffledtabular10M<n<100M0 likes279 downloads9mo agoHugging Face29carsondial /arctic-full-combined-shuffledtext1M<n<10M0 likes268 downloads11mo agoHugging Face30evinsi /fineweb-edu-Llama-3.2-Instruct-Shuffledtext1M<n<10M0 likes259 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.