occiglot/occiglot-fineweb-v0.5
Occiglot Fineweb v0.5 We present a preliminary version of the multilingual Occiglot Fineweb corpus. In this early form, the dataset contains roughly 230M heavily cleaned documents from 10 languages. Occiglot Fineweb builds on our existing collection of curated datasets and pre-filtered web data. Subsequently, all documents were filtered with language-specific derivatives of the fine-web processing pipeline and globally depuplicated. We are actively working on extending this… See the full description on the dataset page: https://huggingface.co/datasets/occiglot/occiglot-fineweb-v0.5.
1514
No card is published for this repository, or it could not be fetched from Hugging Face right now.
