CoolFace
Datasetpublic

KeisukeMiyamoto/CleanedFineWeb2Edu-jp

CleanedFineWeb2Edu-jp CleanedFineWeb2Edu-jp is a cleaned Japanese web text dataset. This dataset was created from the sample_10BT subset of hotchpotch/fineweb-2-edu-japanese. The source text was refined with MK0727/corpus-refiner-jp. Purpose The main purpose of this dataset is to provide cleaner Japanese web text for language model pretraining and continued pretraining. This dataset keeps Japanese web documents from FineWeb2-Edu while reducing boilerplate… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/CleanedFineWeb2Edu-jp.

sourceHugging Faceodc-byupdated 1mo agoView on Hugging Face
1likes362downloads
Dataset Card

CleanedFineWeb2Edu-jp

CleanedFineWeb2Edu-jp is a cleaned Japanese web text dataset.

This dataset was created from the sample_10BT subset of `hotchpotch/fineweb-2-edu-japanese`. The source text was refined with `MK0727/corpus-refiner-jp`.

Purpose

The main purpose of this dataset is to provide cleaner Japanese web text for language model pretraining and continued pretraining.

This dataset keeps Japanese web documents from FineWeb2-Edu while reducing boilerplate, navigation fragments, repeated links, metadata-like text, and other low-value lines. It is intended for use cases that need broad Japanese web text with less page-structure noise than the original subset.

Dataset Creation

After applying MK0727/corpus-refiner-jp, documents were filtered out when the refined text was 200 characters or shorter, or when it was 4096 tokens or longer with sbintuitions/modernbert-ja-130m.

Dataset Statistics

ItemValue
Source sizeabout 12.8B Gemma 4 tokens
Final sizeabout 9.95B Gemma 4 tokens
Retention rateabout 78%

Data Fields

FieldDescription
textCleaned Japanese text
idOriginal document ID
dumpCommon Crawl dump name
urlSource URL
dateCrawl date
file_pathOriginal Common Crawl file path
output_charCharacter count of text
output_tokenToken count of text using google/gemma-4-26B-A4B-it

License

This dataset is released under the ODC Attribution License (odc-by).


Support Lambda

Lambda is an open-source project for building small Japanese language models from scratch. As a student, I have funded this project with income from my part-time job, but the growing training costs are becoming difficult to cover.

Your support helps cover GPU costs and develop larger models. Thank you for helping Lambda continue to grow.

Vast.ai

Vast.ai offers affordable cloud GPUs for AI training, with NVIDIA H100 SXM GPUs available from around $1.54 per hour. If you purchase credits through the link below, I receive 3% in GPU credits at no extra cost to you.

https://cloud.vast.ai/?ref_id=521936

Ko-fi

Support Lambda with a donation starting from $5.

<a href="https://ko-fi.com/lambdallm"> <img src="assets/supportmeonkofibadgeblue.png" alt="Support Lambda on Ko-fi" width="240"> </a>