CoolFace
Datasetpublic

HuggingFaceFW/fineweb

๐Ÿท FineWeb 15 trillion tokens of the finest data the ๐ŸŒ web has to offer What is it? The ๐Ÿท FineWeb dataset consists of more than 18.5T tokens (originally 15T tokens) of cleaned and deduplicated english web data from CommonCrawl. The data processing pipeline is optimized for LLM performance and ran on the ๐Ÿญ datatrove library, our large scale data processing library. ๐Ÿท FineWeb was originally meant to be a fully open replication of ๐Ÿฆ… RefinedWeb, with aโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb.

sourceHugging Faceodc-byupdated 1y agoView on Hugging Face
3.4klikes389kdownloads
settings

This repository belongs to HuggingFaceFW on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namefineweb
visibilitypublic
licenceodc-by
gatedno
ownerHuggingFaceFW
Account settings