CoolFace
Datasetpublic

nizarun/FineWeb-Edu-Arabic-24M

English العربية FineWeb-Edu Arabic 24M An Arabic-only pretraining corpus of 24,794,425 complete documents, translated from the sample-350BT configuration of FineWeb-Edu. It contains 34.86 billion Arabic tokenizer tokens and preserves the original FineWeb-Edu document IDs, source scores, and detailed translation diagnostics. Highlight Saudi architecture shaped by place. A well-translated tour of how builders in Najd, the Gulf coast, Hejaz, and Asir adapted local… See the full description on the dataset page: https://huggingface.co/datasets/nizarun/FineWeb-Edu-Arabic-24M.

sourceHugging Faceodc-byupdated 28d agoView on Hugging Face
0likes464downloads
settings

This repository belongs to nizarun on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameFineWeb-Edu-Arabic-24M
visibilitypublic
licenceodc-by
gatedno
ownernizarun
Account settings
nizarun/FineWeb-Edu-Arabic-24M · CoolFace