CoolFace
Datasetpublic

mariklolik/AraToken-FineWeb2-HQ-ar

AraToken FineWeb2-HQ Arabic splits These are the exact document splits used in AraToken (code). They were drawn from the arb_Arab part of epfml/FineWeb2-HQ and are redistributed under its ODC-By 1.0 license. A document goes to a split by blake2b(id, digest_size=8) mod 1000, so the splits are disjoint by construction. split buckets role documents characters lm-train 0–599 LEP and CPT adaptation 662,685 2.00 B tok-train 600–899 tokenizer training and pruning 159,826… See the full description on the dataset page: https://huggingface.co/datasets/mariklolik/AraToken-FineWeb2-HQ-ar.

sourceHugging Faceodc-byupdated 4d agoView on Hugging Face
0likes56downloads
3 commits on main
94dfedf4d ago

Add dataset card

mariklolik
f3a29554d ago

Add files using upload-large-folder tool

mariklolik
91152974d ago

initial commit

mariklolik