CoolFace
Datasetpublic

hotchpotch/fineweb-2-edu-japanese

šŸ· FineWeb2 Edu Japanese: High-Quality Educational Japanese Dataset This dataset consists of 120 million texts (approximately 89.3B tokens) filtered from the 376 million Japanese texts in FineWeb2 that were deemed educational. The following subsets are also provided: default: Approximately 120M texts (120 million texts) totaling around 89.3B tokens sample_10BT: A random sample of about 10B tokens from the default dataset small_tokens: Data composed solely of texts with 512… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/fineweb-2-edu-japanese.

sourceHugging Faceodc-byupdated 1y agoView on Hugging Face
34likes3.2kdownloads
settings

This repository belongs to hotchpotch on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namefineweb-2-edu-japanese
visibilitypublic
licenceodc-by
gatedno
ownerhotchpotch
Account settings
hotchpotch/fineweb-2-edu-japanese Ā· CoolFace