lbourdois/fineweb-2-trimming
Description Version of FineWeb2 where only 124 languages were kept.For each of them we kept the first 200,000 texts (less if there are not as many available for a given language). The purpose of this dataset is to offer a light version (only 44GB against 8.67 TB for the original dataset) in order to be able to trim models. For more information on the trimming method, we invite you to consult this blog post. Citations FineWeb-2… See the full description on the dataset page: https://huggingface.co/datasets/lbourdois/fineweb-2-trimming.
Update README.md
Update README.md
Update README.md
Update README.md
Upload dataset
Update README.md
Update README.md
Update README.md
Upload dataset
Update README.md
Upload dataset
Update README.md
Upload dataset
Update README.md
Upload dataset
Update README.md
Upload dataset
Update README.md
Update README.md
Delete new_Deva
Update README.md
Upload dataset
Update README.md
Upload dataset
Update README.md
Upload dataset
Update README.md
Upload dataset
Update README.md
Update README.md
Upload dataset
Update README.md
Upload dataset
Delete pms
Update README.md
Upload dataset
Update README.md
Update README.md
Upload dataset
Update README.md
Upload dataset
Update README.md
Upload dataset
Delete lin
Delete lim
Update README.md
Upload dataset
Update README.md
Upload dataset
Update README.md
