sail/regmix-data-sample
RegMix Data Sample Dataset Description The RegMix Data Sample is a curated dataset derived from the Pile-Uncopyrighted, specifically designed for the RegMix paper (https://huggingface.co/papers/2407.01492). This dataset aims to facilitate the automatic identification of high-performing data mixtures for language model pre-training by formulating it as a regression task. Key Features: Size: Approximately 20GB disk space, 5B tokens Distribution:… See the full description on the dataset page: https://huggingface.co/datasets/sail/regmix-data-sample.
2421
Update README.md
Update README.md
Update README.md
Create README.md
update sampled files
initial commit
