CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01adedejimakinde /yoruba-normalization-pairs Normalization pairs dataset What this is 24,475 pairs of Yorùbá text, each a corrupted form next to its canonical form, labelled by corruption type. I built it for testing orthographic normalization code. The library This dataset was built alongside yotext, a Python library for Yorùbá orthographic normalization and diacritic restoration. The library is on PyPI at https://pypi.org/project/yotext/ and the source is at… See the full description on the dataset page: https://huggingface.co/datasets/adedejimakinde/yoruba-normalization-pairs.texttext-generation10K<n<100K1 likes100 downloads11d agoHugging Face02skypro1111 /uk-text-normalization Український TTS-нормалізатор — датасет Пари «письмовий текст → як його вимовляють» для української. Числа, дати, час, гроші, одиниці, скорочення, коди, телефони, IBAN, домени, пошта, римські цифри, латинські вкраплення — те, що треба розгорнути словами перед синтезом мовлення. {"task_id": 0, "combo_names": ["Кількісні числівники (написані цифрами)", "Порядкові числівники (написані цифрами з закінченням)"], "original": "На 1-й полиці стоять 4 книги."… See the full description on the dataset page: https://huggingface.co/datasets/skypro1111/uk-text-normalization.texttext-generation1K<n<10K2 likes77 downloads1mo agoHugging Face03zmsali /bangla-dialect-normalization Bangla Dialect Normalization Dataset A parallel corpus mapping standard Bangla to five regional Bangla dialects, built from the Vashantor dataset. Each row contains the same sentence in standard Bangla and Banglish (romanized), alongside its dialect Bangla and dialect Banglish equivalent, plus an English gloss. Regions covered Barishal, Chittagong, Mymensingh, Noakhali, Sylhet Schema Field Description standard_bangla Sentence in standard… See the full description on the dataset page: https://huggingface.co/datasets/zmsali/bangla-dialect-normalization.texttranslation10K<n<100K0 likes65 downloads22d agoHugging Face04authormist /config-normalization-0909Controlled path-normalization probe. No third-party data. textn<1K0 likes64 downloads12d agoHugging Face05razhan /script_normalization_ckb Script normalization CKB — noisy → standard Sorani (script_normalization_ckb) Nearly 6M sentence pairs: the text column holds Central Kurdish written with non-standard or distorted characters, summary holds the same sentence in standard Sorani orthography. Use it for text normalisation, spell correction, or to learn the character-level mapping rules. At a glance Rows 5,997,025 — train 5,330,689 / test 666,336 Columns text (noisy input), summary… See the full description on the dataset page: https://huggingface.co/datasets/razhan/script_normalization_ckb.text1M<n<10M0 likes45 downloads1d agoHugging Face06Zarinaaa /kyrgyz-text-normalization Kyrgyz Text Normalization Dataset A dataset for training and evaluating Kyrgyz text normalization systems. Released subset accompanying "Kyrgyz Text Normalization: A Comparative Study of Neural and Rule-Based Approaches" (MeLLM Workshop @ ACL 2026). What is in this release This is a representative 20,000-pair subset of a larger 1.67M-pair training corpus, plus the full 1,000-example human-verified test set used in the paper. Split Examples Source Verification… See the full description on the dataset page: https://huggingface.co/datasets/Zarinaaa/kyrgyz-text-normalization.texttext-generation10K<n<100K0 likes39 downloads4mo agoHugging Face07mschonhardt /georges-1913-normalization Normalized Georges 1913 Description This dataset was created as part of the Burchard's Dekret Digital project (www.burchards-dekret-digital.de), funded by the Academy of Sciences and Literature | Mainz. It is based on 55,000 lemmata from Karl Georges, Ausführliches lateinisch-deutsches Handwörterbuch, Hannover 1913 (Georges 1913) and was developed to train models for normalization tasks in the context of medieval Latin. The dataset consists of approximately 5 million… See the full description on the dataset page: https://huggingface.co/datasets/mschonhardt/georges-1913-normalization.text1M<n<10M0 likes37 downloads2y agoHugging Face08sellersew /carrot-engine-normalization-translation-v2text10M<n<100M1 likes35 downloads3y agoHugging Face09yagmurtuncer /turkish-chat-normalization-mini Turkish Chat Normalization Mini turkish-chat-normalization-mini is a web-derived and rule-degraded Turkish text normalization dataset designed for rewriting noisy, informal, unpunctuated, or diacritics-missing Turkish text into cleaner and more readable Turkish. The dataset does not contain private user messages, chat logs, social media comments, complaint records, or scraped personal conversations. Source sentences are collected from open Turkish web resources, while the input… See the full description on the dataset page: https://huggingface.co/datasets/yagmurtuncer/turkish-chat-normalization-mini.texttext-generation10K<n<100K0 likes34 downloads4mo agoHugging Face10yagmurtuncer /turkish-text-normalization 🇹🇷 Turkish Text Normalization (TN / ITN) A deterministic, rule-based dataset of Turkish written ↔ spoken pairs for Text Normalization (TN) and Inverse Text Normalization (ITN) — mapping digit/symbol forms (1.500 TL, %25, 15.07.2026) to their fully spoken Turkish words (bin beş yüz lira, yüzde yirmi beş, on beş temmuz iki bin yirmi altı) and back. This is a common, high-value preprocessing step for Turkish ASR post-processing and TTS front-ends, where numbers, dates, currencies… See the full description on the dataset page: https://huggingface.co/datasets/yagmurtuncer/turkish-text-normalization.texttext-generation10K<n<100K0 likes30 downloads2mo agoHugging Face11larrylawl /chinese-lexical-normalization chinese-lexical-normalization This dataset contains informal-formal-explanation triples from the chinese-lexical-normalization dataset. Note that there are duplicate informal-formal pairs due to multiple explanations. Example usage: from datasets import load_dataset dataset = load_dataset("larrylawl/chinese-lexical-normalization") text1K<n<10K0 likes29 downloads3y agoHugging Face12llmsql-bench /prompts_for_tables_normalization_and_new_sqlstext100K<n<1M0 likes28 downloads4mo agoHugging Face13dataautogpt3 /normalization_faces Dataset Card for "normalization_faces" More Information needed imagen<1K4 likes26 downloads3y agoHugging Face14kenenbek /gemma-russian-normalization-datasettext100K<n<1M0 likes26 downloads11mo agoHugging Face15GoktugD /turkish-text-normalization-1m Turkish Text Normalization 1M v2 Kontrollü altı gürültü türüyle Türkçe metin normalizasyon çiftleri. Doğrulanmış boyut Train: 980,000 Validation: 10,000 Test: 10,000 Toplam: 1,000,000 Ana görev sütunları: id, noisy_text, normalized_text, noise_type Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type, provenance… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-text-normalization-1m.texttext-generation1M<n<10M0 likes24 downloads1mo agoHugging Face16jaio98 /basque_dialect_normalizationtext1K<n<10K0 likes23 downloads4mo agoHugging Face17LegionIntel /date_string_normalizationtext10K<n<100K0 likes20 downloads2y agoHugging Face18distilabel-internal-testing /deita-no-normalization Dataset Card for deita-no-normalization This dataset has been created with Distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/distilabel-internal-testing/deita-no-normalization/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/distilabel-internal-testing/deita-no-normalization.tabular1K<n<10K0 likes19 downloads2y agoHugging Face19Yasshhhh /adaption-telugu-normalization This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-telugu_normalization This dataset contains a collection of user queries and statements written in Telugu, covering diverse topics such as mobile troubleshooting, cyber security threats, insurance renewals, and emergency services. The samples vary from short keywords to detailed problem descriptions involving scams, natural disasters, and administrative procedures. It represents… See the full description on the dataset page: https://huggingface.co/datasets/Yasshhhh/adaption-telugu-normalization.text1K<n<10K0 likes19 downloads3mo agoHugging Face20psidharth567 /normalizationtext100K<n<1M0 likes18 downloads11mo agoHugging Face21DigitalUmuganda /Text_Normalization_Challenge_Unittests_Eng_Fratextn<1K0 likes17 downloads11mo agoHugging Face22GoktugD /turkish-datetime-normalization-500k Turkish Datetime Normalization 500K v2 Türkçe tarih-saat ifadelerini ISO-8601 ve Europe/Istanbul saat dilimine eşler. Doğrulanmış boyut Train: 490,000 Validation: 5,000 Test: 5,000 Toplam: 500,000 Ana görev sütunları: id, text, normalized_datetime, timezone Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type, provenance… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-datetime-normalization-500k.texttext-generation100K<n<1M0 likes17 downloads1mo agoHugging Face23djelia /bm-text-normalizationgated bm-text-normalization Bambara (Bamanankan) orthographic normalisation: map a non-standard spelling to its standard form. 4,877 short phrase-level pairs in a single config, bamadaba. Load from datasets import load_dataset train = load_dataset("djelia/bm-text-normalization", "bamadaba", split="train") dev = load_dataset("djelia/bm-text-normalization", "bamadaba", split="dev") test = load_dataset("djelia/bm-text-normalization", "bamadaba", split="test") # rows… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bm-text-normalization.texttext-generation1K<n<10K0 likes13 downloads2mo agoHugging Face24mengdili /Marco-train-K-16-alpha-2-k-8-type-log-clipping-False-normalization-Falsetext10K<n<100K0 likes12 downloads1y agoHugging Face25userdavek /amharic_summarization_mLongT5_predictions_no_normalization_appliedtext1K<n<10K0 likes12 downloads9mo agoHugging Face26DenysKovalML /ukrainian-dialect-normalizationtext100K<n<1M0 likes12 downloads5mo agoHugging Face27ClarusC64 /clinical-narrative-implicit-normalization-bias-v0.4 Implicit Normalization Bias Clinical Narrative Integrity v0.4 Purpose This dataset tests whether a model: Avoids assuming normality when data is missing Resists default reassurance Preserves honest narrative boundaries Treats “normal” as a claim, not a default You are measuring baseline discipline. Why this dataset exists Clinical notes often omit information. A failure mode distinct from hallucinated negatives is more subtle: Turning… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-narrative-implicit-normalization-bias-v0.4.tabularn<1K0 likes11 downloads8mo agoHugging Face28LRAI /task-normalization-chip2020 Dataset Card for "task-normalization-chip2020" More Information needed text10K<n<100K0 likes10 downloads3y agoHugging Face29lunahr /normalization-data-mixed Normalization Dataset (Mixed) This dataset is a collection of 50000 rows originating from various sources: Wikipedia - 20000 rows PersonaChat truecased - 20000 rows Synthetic edge case data - 5000 rows Synthetic quoted text data - 5000 rows The synthetic data has been generated using GPT-5.3 models. The other data was sourced from the original Hugging Face sources. This dataset can be used to train text normalizers that convert badly formatted English into correct English.… See the full description on the dataset page: https://huggingface.co/datasets/lunahr/normalization-data-mixed.text10K<n<100K0 likes10 downloads2mo agoHugging Face30karanverma19 /Advanced_CodeMix_Normalization_Dataset_India Evaluation & Benchmarking To validate dataset usefulness, normalization accuracy can be evaluated using: Exact Match Accuracy BLEU Score for text similarity Human evaluation for real-world correctness This dataset is designed to improve performance of multilingual NLP systems in handling noisy, code-mixed Indian queries. Data Transformation Approach The dataset was created by transforming real-world code-mixed queries into structured English. Variations include:… See the full description on the dataset page: https://huggingface.co/datasets/karanverma19/Advanced_CodeMix_Normalization_Dataset_India.textn<1K0 likes10 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.