CoolFace
15 results

shuffled

HuggingFaceFW /fineweb_edu_100BT-shuffled FineWeb-Edu 100BT (Shuffled) A globally shuffled version of HuggingFaceFW/fineweb_edu_100BT. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B tokens as fineweb_edu_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining. How It Was Created The unshuffled dataset was loaded into memory, shuffled… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb_edu_100BT-shuffled.tabular100M<n<1B6 likes4.3k downloads7mo agoHugging Facelaion /laions-got-talent-shuffled-with-long-captions1 likes2.3k downloads1y agoHugging FaceZephyr271828 /dclm-llama3-tokenized-shuffled0 likes2.1k downloads5mo agoHugging FaceDanielGallagherIRE /FineWeb-Edu-10B-Shuffledtext1M<n<10M0 likes2k downloads3mo agoHugging FaceMuesli1 /dclm-baseline-1.0-llama3-tokenized-shuffled !! Note: this dataset is currently being uploaded and processed. The .bin files are intermediate files to allow shuffling. !! DCLM-Baseline Pretokenized (LLaMA 3.1, 8192 context) This dataset is a pretokenized and globally shuffled version of DCLM-Baseline (mlfoundations/dclm-baseline-1.0), prepared for large-scale language model pretraining. It is intended to be used as a direct drop-in pretraining corpus for LLaMA 3.1 style training pipelines. The original DCLM-Baseline… See the full description on the dataset page: https://huggingface.co/datasets/Muesli1/dclm-baseline-1.0-llama3-tokenized-shuffled.text0 likes1.8k downloads5mo agoHugging FaceHuggingFaceFW /finepdfs_50BT-dclm_30BT-fineweb_edu_20BT-shuffled FinePDFs 50BT + DCLM 30BT + FineWeb-Edu 20BT (Shuffled) A globally shuffled version of HuggingFaceFW/finepdfs_50BT-dclm_30BT-fineweb_edu_20BT. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B token mixture (50B FinePDFs + 30B DCLM + 20B FineWeb-Edu) but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining —… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_50BT-dclm_30BT-fineweb_edu_20BT-shuffled.text10M<n<100M6 likes1.4k downloads7mo agoHugging Face