CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Scicom-intl /Multilingual-Normalizer Multilingual TTS text normalizer (written → spoken) Training pairs for fine-tuning a small LLM as a text-to-speech normalizer: text is a sentence the way people type it (digits, currency symbols, dates, phone numbers, …) and normalized is the exact spoken form, in the same language, with nothing left that a TTS model cannot say. 52,698 rows — 16 monolingual locales and 6 Malaysian code-switched pairs. Every row is digit-free on the spoken side. from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Multilingual-Normalizer.texttext-generation100K<n<1M0 likes205 downloads18d agoHugging Face02adedejimakinde /yoruba-normalization-pairs Normalization pairs dataset What this is 24,475 pairs of Yorùbá text, each a corrupted form next to its canonical form, labelled by corruption type. I built it for testing orthographic normalization code. The library This dataset was built alongside yotext, a Python library for Yorùbá orthographic normalization and diacritic restoration. The library is on PyPI at https://pypi.org/project/yotext/ and the source is at… See the full description on the dataset page: https://huggingface.co/datasets/adedejimakinde/yoruba-normalization-pairs.texttext-generation10K<n<100K1 likes101 downloads12d agoHugging Face03francescortu /DistillDetect-normalized-traces DistillDetect — format-normalized teacher traces Teacher responses from Reference-Based Distillation Detection in LLMs (arXiv:2607.09692), rewritten so that every teacher uses the same output format. 7,918 rows across 8 teacher/prompt-set pairs. Why this exists In the released data each teacher emits a structurally different response, so a student trained on it — and any detector trained to attribute it — can key on surface format instead of the teacher's actual… See the full description on the dataset page: https://huggingface.co/datasets/francescortu/DistillDetect-normalized-traces.texttext-generation1K<n<10K0 likes78 downloads29d agoHugging Face04skypro1111 /uk-text-normalization Український TTS-нормалізатор — датасет Пари «письмовий текст → як його вимовляють» для української. Числа, дати, час, гроші, одиниці, скорочення, коди, телефони, IBAN, домени, пошта, римські цифри, латинські вкраплення — те, що треба розгорнути словами перед синтезом мовлення. {"task_id": 0, "combo_names": ["Кількісні числівники (написані цифрами)", "Порядкові числівники (написані цифрами з закінченням)"], "original": "На 1-й полиці стоять 4 книги."… See the full description on the dataset page: https://huggingface.co/datasets/skypro1111/uk-text-normalization.texttext-generation1K<n<10K2 likes63 downloads1mo agoHugging Face05atrevidasadia /dair-ai-emotion-normalized-instruction-input-output dair-ai emotion | normalized Summary Dataset ID: 143 Type: normalized Rows: 16,000 Source: dair-ai/emotion Dataset Sources #143 dair-ai emotion | normalized [normalized | 16,000 rows] Notes Edited and Exported from the Kitsune Training Suite (Forge) Review the dataset artifact and metadata before publishing. Citation > via dair-ai @inproceedings{saravia-etal-2018-carer, title = "{CARER}: Contextualized Affect… See the full description on the dataset page: https://huggingface.co/datasets/atrevidasadia/dair-ai-emotion-normalized-instruction-input-output.texttext-generation10K<n<100K0 likes44 downloads29d agoHugging Face06Zarinaaa /kyrgyz-text-normalization Kyrgyz Text Normalization Dataset A dataset for training and evaluating Kyrgyz text normalization systems. Released subset accompanying "Kyrgyz Text Normalization: A Comparative Study of Neural and Rule-Based Approaches" (MeLLM Workshop @ ACL 2026). What is in this release This is a representative 20,000-pair subset of a larger 1.67M-pair training corpus, plus the full 1,000-example human-verified test set used in the paper. Split Examples Source Verification… See the full description on the dataset page: https://huggingface.co/datasets/Zarinaaa/kyrgyz-text-normalization.texttext-generation10K<n<100K0 likes32 downloads4mo agoHugging Face07deltakitsune /dair-ai-emotion-normalized-instruction-input-output dair-ai emotion | normalized Summary Dataset ID: 143 Type: normalized Rows: 16,000 Source: dair-ai/emotion Dataset Sources #143 dair-ai emotion | normalized [normalized | 16,000 rows] Notes Edited and Exported from the Kitsune Training Suite (Forge) Review the dataset artifact and metadata before publishing. Citation > via dair-ai @inproceedings{saravia-etal-2018-carer, title = "{CARER}: Contextualized Affect Representations for… See the full description on the dataset page: https://huggingface.co/datasets/deltakitsune/dair-ai-emotion-normalized-instruction-input-output.texttext-generation10K<n<100K0 likes26 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.