jiviteshjn/fineweb-edu-zh-chengyu-cpt
Fineweb-Edu Chinese — Chengyu-Tagged Continued-Pretraining Corpus A 3.74M-document Chinese corpus (~7.8B tokens) for continued pretraining on cultural knowledge in figurative language, built from the highest-quality tier of opencsg/Fineweb-Edu-Chinese-V2.1. Each document is educational Chinese text containing at least one culturally vetted chengyu, with an appended 【成语注释】 knowledge block listing every matched idiom's figurative meaning(s) and classical source citation. This is a… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/fineweb-edu-zh-chengyu-cpt.
Fineweb-Edu Chinese — Chengyu-Tagged Continued-Pretraining Corpus
A 3.74M-document Chinese corpus (~7.8B tokens) for continued pretraining on cultural knowledge in figurative language, built from the highest-quality tier of opencsg/Fineweb-Edu-Chinese-V2.1. Each document is educational Chinese text containing at least one culturally vetted chengyu, with an appended 【成语注释】 knowledge block listing every matched idiom's figurative meaning(s) and classical source citation.
This is a separate dataset from our mC4-based Chinese corpus (same idiom list and pipeline, different base corpus — not merged, and deduplication was performed only within this corpus).
Key numbers
Document format
data/tagged_*.json.gz — one JSON object per line:
{
"text": "<document text>\n\n【成语注释】\n<idiom>:<figurative meaning(s)>(出处/字面:<classical source>)\n...",
"source": "fineweb-edu-zh-4_5",
"orig_source": "<upstream source label, e.g. CCI3, map-cc>",
"score": <upstream educational-quality score>,
"shard": 0, "line": 123,
"matched_idioms": ["..."],
"url": "", "timestamp": ""
}How it was built
- Idiom vetting — a merged 31K-entry chengyu dataset (chinese-xinhua, chengyudata, FuxiBench, IdiomKB) filtered by an LLM judge on cultural/figurative content (not commonness): 27,297 idioms kept. Per-entry verdicts: `idioms/filterdecisionszh.jsonl`; the vetted list: `idioms/idiomsmergedllmformattedfigurativeonly.jsonl`.
- Extraction — the full 4_5 (highest) quality tier of Fineweb-Edu-Chinese-V2.1 (9,745 parquet files, 17.79M docs, ~46B tokens) scanned with Aho-Corasick multi-pattern matching. Educational Chinese is remarkably chengyu-dense: 47.2% of docs matched (vs ~30% for mC4 web text). Quality gates: 150–100,000 chars and a word-list gate (>20 distinct idioms) that dropped 181K docs — largely chengyu study lists, which would otherwise pollute training.
- Deduplication — exact (blake2b) + MinHash-LSH near-dedup (128 perms, ~0.75 Jaccard): 52,731 exact + 2,225,650 near duplicates removed (27% of candidates — educational articles are heavily reposted across sites).
- Frequency-capped selection (rare-idiom-first) — idioms in ascending candidate count each top up to 10,000 documents; selecting a document credits all its idioms, so common idioms fill via co-occurrence and rare ones never lose out.
- Tagging — the
【成语注释】knowledge block: up to 2 deduplicated figurative meanings + the first literal meaning / classical source citation per matched idiom.
Usage
from datasets import load_dataset
ds = load_dataset("jiviteshjn/fineweb-edu-zh-chengyu-cpt", data_files="data/*.json.gz", split="train")hf download jiviteshjn/fineweb-edu-zh-chengyu-cpt --repo-type dataset --local-dir fwe_zh_chengyu_cptDocuments average ~2.1K tokens; sequence packing recommended for training.
Provenance & licenses
- Base corpus: opencsg/Fineweb-Edu-Chinese-V2.1, Apache 2.0. Note the upstream card states commercial applications require written permission from OpenCSG (lorraineg@opencsg.com) — this derivative inherits that condition; it is released for research use.
- Idiom meanings merged from public chengyu dictionaries.
- The upstream
score(educational quality) andorig_sourcefields are preserved per document for downstream filtering.
Compared to our mC4-based Chinese corpus: cleaner, denser educational prose with a higher idiom rate, but lower rare-idiom coverage (20,440 vs 26,727 of 27,297) — curated educational content surfaces the literary chengyu tail less than broad web crawl does. The two corpora are complementary.
