CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01skypro1111 /uk-text-normalization Український TTS-нормалізатор — датасет Пари «письмовий текст → як його вимовляють» для української. Числа, дати, час, гроші, одиниці, скорочення, коди, телефони, IBAN, домени, пошта, римські цифри, латинські вкраплення — те, що треба розгорнути словами перед синтезом мовлення. {"task_id": 0, "combo_names": ["Кількісні числівники (написані цифрами)", "Порядкові числівники (написані цифрами з закінченням)"], "original": "На 1-й полиці стоять 4 книги."… See the full description on the dataset page: https://huggingface.co/datasets/skypro1111/uk-text-normalization.texttext-generation1K<n<10K2 likes77 downloads1mo agoHugging Face02Zarinaaa /kyrgyz-text-normalization Kyrgyz Text Normalization Dataset A dataset for training and evaluating Kyrgyz text normalization systems. Released subset accompanying "Kyrgyz Text Normalization: A Comparative Study of Neural and Rule-Based Approaches" (MeLLM Workshop @ ACL 2026). What is in this release This is a representative 20,000-pair subset of a larger 1.67M-pair training corpus, plus the full 1,000-example human-verified test set used in the paper. Split Examples Source Verification… See the full description on the dataset page: https://huggingface.co/datasets/Zarinaaa/kyrgyz-text-normalization.texttext-generation10K<n<100K0 likes39 downloads4mo agoHugging Face03pavanBuduguppa /asr_inverse_text_normalization1 likes31 downloads4y agoHugging Face04yagmurtuncer /turkish-text-normalization 🇹🇷 Turkish Text Normalization (TN / ITN) A deterministic, rule-based dataset of Turkish written ↔ spoken pairs for Text Normalization (TN) and Inverse Text Normalization (ITN) — mapping digit/symbol forms (1.500 TL, %25, 15.07.2026) to their fully spoken Turkish words (bin beş yüz lira, yüzde yirmi beş, on beş temmuz iki bin yirmi altı) and back. This is a common, high-value preprocessing step for Turkish ASR post-processing and TTS front-ends, where numbers, dates, currencies… See the full description on the dataset page: https://huggingface.co/datasets/yagmurtuncer/turkish-text-normalization.texttext-generation10K<n<100K0 likes30 downloads2mo agoHugging Face05GoktugD /turkish-text-normalization-1m Turkish Text Normalization 1M v2 Kontrollü altı gürültü türüyle Türkçe metin normalizasyon çiftleri. Doğrulanmış boyut Train: 980,000 Validation: 10,000 Test: 10,000 Toplam: 1,000,000 Ana görev sütunları: id, noisy_text, normalized_text, noise_type Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type, provenance… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-text-normalization-1m.texttext-generation1M<n<10M0 likes24 downloads1mo agoHugging Face06DigitalUmuganda /Text_Normalization_Challenge_Unittests_Eng_Fratextn<1K0 likes17 downloads11mo agoHugging Face07djelia /bm-text-normalizationgated bm-text-normalization Bambara (Bamanankan) orthographic normalisation: map a non-standard spelling to its standard form. 4,877 short phrase-level pairs in a single config, bamadaba. Load from datasets import load_dataset train = load_dataset("djelia/bm-text-normalization", "bamadaba", split="train") dev = load_dataset("djelia/bm-text-normalization", "bamadaba", split="dev") test = load_dataset("djelia/bm-text-normalization", "bamadaba", split="test") # rows… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bm-text-normalization.texttext-generation1K<n<10K0 likes13 downloads2mo agoHugging Face08djelia /text-normalization-benchmarkgated text-normalization-benchmark The raw Argilla 2.8.0 export of a Bambara (Bamanankan) text-normalization project: 160 records from four in-house corpora, each with the annotator's standard-orthography rewrite. 96 carry a submitted response; 64 were discarded. For a ready-to-score evaluation set, use djelia/bm-text-normalization-benchmark, the cleaned export of the 96 finished annotations. The repo is gated: request access on the Hub and run hf auth login. Load from… See the full description on the dataset page: https://huggingface.co/datasets/djelia/text-normalization-benchmark.texttext-generationn<1K0 likes9 downloads2mo agoHugging Face09Saikrishna2403 /ml_ta_text_normalizationtext10K<n<100K0 likes8 downloads2y agoHugging Face10Saikrishna2403 /tamil_ml_text_normalizationtext10K<n<100K0 likes7 downloads2y agoHugging Face11djelia /bm-text-normalization-benchmarkgated bm-text-normalization-benchmark A small human-annotated evaluation set for Bambara (Bamanankan) orthographic normalisation: 96 real-world Bambara strings, each paired with a hand-written standard-orthography rewrite. It is the cleaned export of the finished annotations from djelia/text-normalization-benchmark. Load from datasets import load_dataset # the current, whitespace-clean evaluation set bench = load_dataset("djelia/bm-text-normalization-benchmark"… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bm-text-normalization-benchmark.texttext-generationn<1K0 likes7 downloads2mo agoHugging Face12Saikrishna2403 /ta_ml_text_normalizationtext10K<n<100K0 likes6 downloads2y agoHugging Face13SPEAK-PP /Inverse_Text_Normalization_Sinhalatext1K<n<10K0 likes6 downloads6mo agoHugging Face14shubham-Bgs /Text-Normalization-Hinditext1K<n<10K1 likes5 downloads2y agoHugging Face15KGSAGAR /Text_Normalization_HindiThe dataset is aimed to train model to take a sentence in Hindi language and normalize specific entities, including: Dates (any format) Currencies Scientific units Hindi input output १९८४ में इंदिरा गांधी की हत्या के बाद उनके बेटे राजीव गांधी कॉंग्रेस के नेता और प्रधानमंत्री बने निन्टीन इटी फोर में इंदिरा गांधी की हत्या के बाद उनके बेटे राजीव गांधी कॉंग्रेस के नेता और प्रधानमंत्री बने प्रथम स्थाई बस्तियों ने 9000 वर्ष पूर्व स्वरुप लिया था प्रथम स्थाई बस्तियों ने नाइन थाउजंड वर्ष… See the full description on the dataset page: https://huggingface.co/datasets/KGSAGAR/Text_Normalization_Hindi.n<1K0 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.