CoolFace
Datasetpublic

Brainquiver/general-web-it-202608

General · Web · Italian · 2026-08 Italian pretraining text, built from the Italian portion of EPFL's FineWeb2-HQ, which is the high quality slice of FineWeb-2. Every document passes one character-level cleaner and a repetition filter. 21,065,052 documents and 66,158,573,443 characters of Italian prose. Contents Config Documents Characters Upstream fineweb2-hq-ita_Latn 21,065,052 66,158,573,443 epfml/FineWeb2-HQ, ita_Latn The character count is… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-it-202608.

sourceHugging Faceodc-byupdated 26d agoView on Hugging Face
0likes554downloads
settings

This repository belongs to Brainquiver on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namegeneral-web-it-202608
visibilitypublic
licenceodc-by
gatedno
ownerBrainquiver
Account settings
Brainquiver/general-web-it-202608 · CoolFace