datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mc4_3.1.0_fi_cleaned
Dataset Card for "mc4_3.1.0_fi_cleaned"
More Information needed
mc4-zh-idiom-cpt
mC4 zh — Idiom-Tagged Continued-Pretraining Corpus
A 9.6M-document Chinese corpus for continued pretraining on cultural knowledge in
figurative language. Each document is natural web text (from the C4/mC4 zh subset)
containing at least one culturally meaningful chengyu, with an appended knowledge
block that lists every matched idiom together with its figurative meaning(s) and
classical source citation.
Built 2026-07-16 as Stage 1 (continue-pretraining data) of the… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/mc4-zh-idiom-cpt.mc4_ja_text_volume_annotatted_dataこのデータセットはmC4の日本語データに対し、文の割合に応じて1~5段階で人手で評価したものです。文の割合はデータセットのscoreに格納しています。
文の割合が20%以下
文の割合が20~40%
文の割合が40~60%
文の割合が60~80%
文の割合が80~100%
データ量は500件と少ないですが、mC4のゴミデータ削除に役立てればと思います。
