datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb-2-sentence-splitFineweb 2 split into sentences. Instances per languages were sampled by us to balance the data w.r.t. Fineweb-edu.
To split the text into sentences we used the sat3-l model from the wtpsplit library.
We fix a sentence threshold of 0.02 and a maximum sentence length of 256.
If you use this dataset, you should cite:
@misc{penedo2025fineweb2pipelinescale,
title={FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language},
author={Guilherme Penedo and… See the full description on the dataset page: https://huggingface.co/datasets/mimir-lcm/fineweb-2-sentence-split.fineweb-edu-350BT-sentence-splitFineweb-edu 350BT subset split into sentences.
To split the text into sentences we used the sat3-l model from the wtpsplit library.
We fix a sentence threshold of 0.02 and a maximum sentence length of 256.
If you use this dataset, you should cite:
@misc{lozhkov2024fineweb-edu,
author = { Lozhkov, Anton and Ben Allal, Loubna and von Werra, Leandro and Wolf, Thomas },
title = { FineWeb-Edu: the Finest Collection of Educational Content },
year = 2024… See the full description on the dataset page: https://huggingface.co/datasets/mimir-lcm/fineweb-edu-350BT-sentence-split.lcm-newtest-lcm
Dataset Card for "test-lcm"
More Information needed
dependency-parsing-lcmlcmds_lcmqjcpbebPreprocessed_fineweb_lcm
