CoolFace
Datasetpublic

jiviteshjn/fineweb-edu-zh-chengyu-cpt

Fineweb-Edu Chinese — Chengyu-Tagged Continued-Pretraining Corpus A 3.74M-document Chinese corpus (~7.8B tokens) for continued pretraining on cultural knowledge in figurative language, built from the highest-quality tier of opencsg/Fineweb-Edu-Chinese-V2.1. Each document is educational Chinese text containing at least one culturally vetted chengyu, with an appended 【成语注释】 knowledge block listing every matched idiom's figurative meaning(s) and classical source citation. This is a… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/fineweb-edu-zh-chengyu-cpt.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
1likes217downloads
settings

This repository belongs to jiviteshjn on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namefineweb-edu-zh-chengyu-cpt
visibilitypublic
licenceapache-2.0
gatedno
ownerjiviteshjn
Account settings
jiviteshjn/fineweb-edu-zh-chengyu-cpt · CoolFace