mimir-lcm/fineweb-2-sentence-split
Fineweb 2 split into sentences. Instances per languages were sampled by us to balance the data w.r.t. Fineweb-edu. To split the text into sentences we used the sat3-l model from the wtpsplit library. We fix a sentence threshold of 0.02 and a maximum sentence length of 256. If you use this dataset, you should cite: @misc{penedo2025fineweb2pipelinescale, title={FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language}, author={Guilherme Penedo and… See the full description on the dataset page: https://huggingface.co/datasets/mimir-lcm/fineweb-2-sentence-split.
Fineweb 2 split into sentences. Instances per languages were sampled by us to balance the data w.r.t. Fineweb-edu.
To split the text into sentences we used the sat3-l model from the wtpsplit library.
We fix a sentence threshold of 0.02 and a maximum sentence length of 256.
If you use this dataset, you should cite:
@misc{penedo2025fineweb2pipelinescale,
title={FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language},
author={Guilherme Penedo and Hynek Kydlíček and Vinko Sabolčec and Bettina Messmer and Negar Foroutan and Amir Hossein Kargaran and Colin Raffel and Martin Jaggi and Leandro Von Werra and Thomas Wolf},
year={2025},
eprint={2506.20920},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2506.20920},
}
@misc{musacchio2026mimirlargescalemultilingualconcept,
title={Mimir: Large-scale Multilingual Concept Modeling},
author={Elio Musacchio and Lucia Siciliani and Pierpaolo Basile},
year={2026},
eprint={2605.25263},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2605.25263},
}