CoolFace
Datasetpublic

mimir-lcm/fineweb-2-sentence-split

Fineweb 2 split into sentences. Instances per languages were sampled by us to balance the data w.r.t. Fineweb-edu. To split the text into sentences we used the sat3-l model from the wtpsplit library. We fix a sentence threshold of 0.02 and a maximum sentence length of 256. If you use this dataset, you should cite: @misc{penedo2025fineweb2pipelinescale, title={FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language}, author={Guilherme Penedo and… See the full description on the dataset page: https://huggingface.co/datasets/mimir-lcm/fineweb-2-sentence-split.

sourceHugging Faceodc-byupdated 4mo agoView on Hugging Face
0likes5.9kdownloads
settings

This repository belongs to mimir-lcm on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namefineweb-2-sentence-split
visibilitypublic
licenceodc-by
gatedno
ownermimir-lcm
Account settings
mimir-lcm/fineweb-2-sentence-split · CoolFace