CoolFace
Datasetpublic

jiviteshjn/fineweb-edu-zh-chengyu-cpt

Fineweb-Edu Chinese — Chengyu-Tagged Continued-Pretraining Corpus A 3.74M-document Chinese corpus (~7.8B tokens) for continued pretraining on cultural knowledge in figurative language, built from the highest-quality tier of opencsg/Fineweb-Edu-Chinese-V2.1. Each document is educational Chinese text containing at least one culturally vetted chengyu, with an appended 【成语注释】 knowledge block listing every matched idiom's figurative meaning(s) and classical source citation. This is a… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/fineweb-edu-zh-chengyu-cpt.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
1likes217downloads
Dataset Card

Fineweb-Edu Chinese — Chengyu-Tagged Continued-Pretraining Corpus

A 3.74M-document Chinese corpus (~7.8B tokens) for continued pretraining on cultural knowledge in figurative language, built from the highest-quality tier of opencsg/Fineweb-Edu-Chinese-V2.1. Each document is educational Chinese text containing at least one culturally vetted chengyu, with an appended 【成语注释】 knowledge block listing every matched idiom's figurative meaning(s) and classical source citation.

This is a separate dataset from our mC4-based Chinese corpus (same idiom list and pipeline, different base corpus — not merged, and deduplication was performed only within this corpus).

Key numbers

Documents3,744,423
Tokens (Qwen3.5 tokenizer, sampled)~7.8B
Idiom annotations18,220,343
Distinct idioms covered20,440 (of a 27,297 vetted list)
Idioms saturated at the 10K/idiom cap479
SourceFineweb-Edu-Chinese-V2.1, 4_5 score tier only (17.79M docs scanned)
Idiom match rate in source46.8% of docs

Document format

data/tagged_*.json.gz — one JSON object per line:

json
{
  "text": "<document text>\n\n【成语注释】\n<idiom>:<figurative meaning(s)>(出处/字面:<classical source>)\n...",
  "source": "fineweb-edu-zh-4_5",
  "orig_source": "<upstream source label, e.g. CCI3, map-cc>",
  "score": <upstream educational-quality score>,
  "shard": 0, "line": 123,
  "matched_idioms": ["..."],
  "url": "", "timestamp": ""
}

How it was built

  1. 1.Idiom vetting — a merged 31K-entry chengyu dataset (chinese-xinhua, chengyudata, FuxiBench, IdiomKB) filtered by an LLM judge on cultural/figurative content (not commonness): 27,297 idioms kept. Per-entry verdicts: `idioms/filterdecisionszh.jsonl`; the vetted list: `idioms/idiomsmergedllmformattedfigurativeonly.jsonl`.
  2. 2.Extraction — the full 4_5 (highest) quality tier of Fineweb-Edu-Chinese-V2.1 (9,745 parquet files, 17.79M docs, ~46B tokens) scanned with Aho-Corasick multi-pattern matching. Educational Chinese is remarkably chengyu-dense: 47.2% of docs matched (vs ~30% for mC4 web text). Quality gates: 150–100,000 chars and a word-list gate (>20 distinct idioms) that dropped 181K docs — largely chengyu study lists, which would otherwise pollute training.
  3. 3.Deduplication — exact (blake2b) + MinHash-LSH near-dedup (128 perms, ~0.75 Jaccard): 52,731 exact + 2,225,650 near duplicates removed (27% of candidates — educational articles are heavily reposted across sites).
  4. 4.Frequency-capped selection (rare-idiom-first) — idioms in ascending candidate count each top up to 10,000 documents; selecting a document credits all its idioms, so common idioms fill via co-occurrence and rare ones never lose out.
  5. 5.Tagging — the 【成语注释】 knowledge block: up to 2 deduplicated figurative meanings + the first literal meaning / classical source citation per matched idiom.

Usage

python
from datasets import load_dataset
ds = load_dataset("jiviteshjn/fineweb-edu-zh-chengyu-cpt", data_files="data/*.json.gz", split="train")
bash
hf download jiviteshjn/fineweb-edu-zh-chengyu-cpt --repo-type dataset --local-dir fwe_zh_chengyu_cpt

Documents average ~2.1K tokens; sequence packing recommended for training.

Provenance & licenses

  • —Base corpus: opencsg/Fineweb-Edu-Chinese-V2.1, Apache 2.0. Note the upstream card states commercial applications require written permission from OpenCSG (lorraineg@opencsg.com) — this derivative inherits that condition; it is released for research use.
  • —Idiom meanings merged from public chengyu dictionaries.
  • —The upstream score (educational quality) and orig_source fields are preserved per document for downstream filtering.

Compared to our mC4-based Chinese corpus: cleaner, denser educational prose with a higher idiom rate, but lower rare-idiom coverage (20,440 vs 26,727 of 27,297) — curated educational content surfaces the literary chengyu tail less than broad web crawl does. The two corpora are complementary.