CoolFace
Datasetpublic

TheFinAI/dolma3_300B_sample_shuffled

dolma3_300B_sample_shuffled Global row-level shuffle of TheFinAI/dolma3_300B_sample. Source data uses per-row Bernoulli sampling (p ≈ 0.0506) from allenai/dolma3_mix-6T-1025-7B to produce ~300B cl100k tokens preserving the original Dolma3 mix ratios. However the source parquets cluster records by sub-source on disk (each ~100K-row parquet groups rows from the same input shard contiguously), which means a small training shuffle buffer would see a non-uniform source mix per… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/dolma3_300B_sample_shuffled.

sourceHugging Faceodc-byupdated 4mo agoView on Hugging Face
0likes46downloads
Dataset Card

dolma3300Bsample_shuffled

Global row-level shuffle of TheFinAI/dolma3_300B_sample.

Source data uses per-row Bernoulli sampling (p ≈ 0.0506) from allenai/dolma3_mix-6T-1025-7B to produce ~300B cl100k tokens preserving the original Dolma3 mix ratios. However the source parquets cluster records by sub-source on disk (each ~100K-row parquet groups rows from the same input shard contiguously), which means a small training shuffle buffer would see a non-uniform source mix per minibatch.

This dataset is a true global shuffle of the source: every row is uniformly randomly assigned to one of 200 output buckets, then each bucket is shuffled in memory. Source mix ratios are unchanged.

Schema: same as source (source, date, text, token_count, category).

Seed: 42.