CoolFace
Datasetpublic

jiviteshjn/fineweb-edu-zh-chengyu-cpt

Fineweb-Edu Chinese — Chengyu-Tagged Continued-Pretraining Corpus A 3.74M-document Chinese corpus (~7.8B tokens) for continued pretraining on cultural knowledge in figurative language, built from the highest-quality tier of opencsg/Fineweb-Edu-Chinese-V2.1. Each document is educational Chinese text containing at least one culturally vetted chengyu, with an appended 【成语注释】 knowledge block listing every matched idiom's figurative meaning(s) and classical source citation. This is a… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/fineweb-edu-zh-chengyu-cpt.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
1likes217downloads

jiviteshjn/fineweb-edu-zh-chengyu-cpt · main · files are served by the source, never re-hosted here