CoolFace
Datasetpublic

TheFinAI/dolma3_300B_sample_shuffled

dolma3_300B_sample_shuffled Global row-level shuffle of TheFinAI/dolma3_300B_sample. Source data uses per-row Bernoulli sampling (p ≈ 0.0506) from allenai/dolma3_mix-6T-1025-7B to produce ~300B cl100k tokens preserving the original Dolma3 mix ratios. However the source parquets cluster records by sub-source on disk (each ~100K-row parquet groups rows from the same input shard contiguously), which means a small training shuffle buffer would see a non-uniform source mix per… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/dolma3_300B_sample_shuffled.

sourceHugging Faceodc-byupdated 4mo agoView on Hugging Face
0likes46downloads

No commit history came back for main. The revision may not exist, or the source declined the request.