CoolFace
Datasetpublicgated

occiglot/occiglot-fineweb-v0.5

Occiglot Fineweb v0.5 We present a preliminary version of the multilingual Occiglot Fineweb corpus. In this early form, the dataset contains roughly 230M heavily cleaned documents from 10 languages. Occiglot Fineweb builds on our existing collection of curated datasets and pre-filtered web data. Subsequently, all documents were filtered with language-specific derivatives of the fine-web processing pipeline and globally depuplicated. We are actively working on extending this… See the full description on the dataset page: https://huggingface.co/datasets/occiglot/occiglot-fineweb-v0.5.

sourceHugging Faceupdated 2y agoView on Hugging Face
15likes14downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.