shuffled
babylm-shuffled-1234-gpt2-smallrun2-babylm-shuffled-2345-gpt2-smallrun2-babylm-shuffled-1234-gpt2-smallSongTonyLi_-_Phi-3.5-mini-instruct-SFT-D_chosen-dpo-mix-shuffled-ggufSongTonyLi_-_Phi-3.5-mini-instruct-CPT-D1_chosen-dpo-mix-shuffled5-ggufSongTonyLi_-_Phi-3.5-mini-instruct-SFT-D_chosen-dpo-mix-shuffled4-ggufalexshengzhili_-_dpo_0908_preference_4_conference_shuffled_2023_checkpoint_30-ggufalexshengzhili_-_phi3-dpo_0908_preference_4_conference_shuffled_2023-gguf
Datasets
All datasets matching “shuffled”fineweb_edu_100BT-shuffled
FineWeb-Edu 100BT (Shuffled)
A globally shuffled version of HuggingFaceFW/fineweb_edu_100BT.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
Dataset Description
This dataset contains the same ~100B tokens as fineweb_edu_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining.
How It Was Created
The unshuffled dataset was loaded into memory, shuffled… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb_edu_100BT-shuffled.laions-got-talent-shuffled-with-long-captionsdclm-llama3-tokenized-shuffledFineWeb-Edu-10B-Shuffleddclm-baseline-1.0-llama3-tokenized-shuffled
!! Note: this dataset is currently being uploaded and processed. The .bin files are intermediate files to allow shuffling. !!
DCLM-Baseline Pretokenized (LLaMA 3.1, 8192 context)
This dataset is a pretokenized and globally shuffled version of DCLM-Baseline (mlfoundations/dclm-baseline-1.0), prepared for large-scale language model pretraining. It is intended to be used as a direct drop-in pretraining corpus for LLaMA 3.1 style training pipelines.
The original DCLM-Baseline… See the full description on the dataset page: https://huggingface.co/datasets/Muesli1/dclm-baseline-1.0-llama3-tokenized-shuffled.finepdfs_50BT-dclm_30BT-fineweb_edu_20BT-shuffled
FinePDFs 50BT + DCLM 30BT + FineWeb-Edu 20BT (Shuffled)
A globally shuffled version of HuggingFaceFW/finepdfs_50BT-dclm_30BT-fineweb_edu_20BT.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
Dataset Description
This dataset contains the same ~100B token mixture (50B FinePDFs + 30B DCLM + 20B FineWeb-Edu) but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining —… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_50BT-dclm_30BT-fineweb_edu_20BT-shuffled.
