datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turkish-diacritics-restoration-1m
Turkish Diacritics Restoration 1M v2
ASCII'ye indirgenmiş Türkçe metinler ve karakterleri geri yüklenmiş hedefleri.
Doğrulanmış boyut
Train: 980,000
Validation: 10,000
Test: 10,000
Toplam: 1,000,000
Ana görev sütunları: id, ascii_text, restored_text
Provenance
Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı
depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type,
provenance… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-diacritics-restoration-1m.diacritics-ro
Romanian Diacritic Restoration Dataset (dexonline)
Data splits for Romanian automatic diacritic restoration (ADR), derived from the dexonline lexicographic corpus.
Dataset Description
Each example contains a input (stripped text without diacritics), a target (gold-standard diacritized text), and an id.
Source: dexonline.ro -- open-access Romanian lexicographic resource
Training pairs: 299,633 examples (NFKD diacritic stripping + cedilla normalization)
Validation: 33,292… See the full description on the dataset page: https://huggingface.co/datasets/klusai/diacritics-ro.
