datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fix_punctuationclassical-chinese-punctuation
Classical Chinese Punctuation Restoration · 文言文断句·标点 💰 (Commercial Dataset)
This is a commercial dataset. A free 200-record sample is provided below
(sample.jsonl); the full 5.3M-pair corpus is available upon request.
📧 To license / purchase or request a quote, email wangeksy@gmail.com.
✅ Cleared for commercial use — built from public-domain classical works.
The task
Restore punctuation and sentence segmentation (句读) to unpunctuated Classical
Chinese — a… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/classical-chinese-punctuation.htest-end-punctuationdhivehi-punctuation-dataset
Dhivehi punctuation restoration dataset
Self-supervised punctuation-restoration data: punctuated Dhivehi text with punctuation stripped and
kept as per-word labels. 695,854 train / 24,896 validation / 25,308 test segments; each segment is a
run of 1–6 consecutive sentences. Splits are by source document.
Sources: alakxender/dhivehi-majlis-transcripts (spoken; minutes appear twice in train with
different segmentations) and 30k articles from alakxender/dhivehi-news-corpus… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-punctuation-dataset.
