CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01adedejimakinde /yoruba-normalization-pairs Normalization pairs dataset What this is 24,475 pairs of Yorùbá text, each a corrupted form next to its canonical form, labelled by corruption type. I built it for testing orthographic normalization code. The library This dataset was built alongside yotext, a Python library for Yorùbá orthographic normalization and diacritic restoration. The library is on PyPI at https://pypi.org/project/yotext/ and the source is at… See the full description on the dataset page: https://huggingface.co/datasets/adedejimakinde/yoruba-normalization-pairs.texttext-generation10K<n<100K1 likes100 downloads11d agoHugging Face02skypro1111 /uk-text-normalization Український TTS-нормалізатор — датасет Пари «письмовий текст → як його вимовляють» для української. Числа, дати, час, гроші, одиниці, скорочення, коди, телефони, IBAN, домени, пошта, римські цифри, латинські вкраплення — те, що треба розгорнути словами перед синтезом мовлення. {"task_id": 0, "combo_names": ["Кількісні числівники (написані цифрами)", "Порядкові числівники (написані цифрами з закінченням)"], "original": "На 1-й полиці стоять 4 книги."… See the full description on the dataset page: https://huggingface.co/datasets/skypro1111/uk-text-normalization.texttext-generation1K<n<10K2 likes77 downloads1mo agoHugging Face03zmsali /bangla-dialect-normalization Bangla Dialect Normalization Dataset A parallel corpus mapping standard Bangla to five regional Bangla dialects, built from the Vashantor dataset. Each row contains the same sentence in standard Bangla and Banglish (romanized), alongside its dialect Bangla and dialect Banglish equivalent, plus an English gloss. Regions covered Barishal, Chittagong, Mymensingh, Noakhali, Sylhet Schema Field Description standard_bangla Sentence in standard… See the full description on the dataset page: https://huggingface.co/datasets/zmsali/bangla-dialect-normalization.texttranslation10K<n<100K0 likes65 downloads22d agoHugging Face04authormist /config-normalization-0909Controlled path-normalization probe. No third-party data. textn<1K0 likes64 downloads12d agoHugging Face05razhan /script_normalization_ckb Script normalization CKB — noisy → standard Sorani (script_normalization_ckb) Nearly 6M sentence pairs: the text column holds Central Kurdish written with non-standard or distorted characters, summary holds the same sentence in standard Sorani orthography. Use it for text normalisation, spell correction, or to learn the character-level mapping rules. At a glance Rows 5,997,025 — train 5,330,689 / test 666,336 Columns text (noisy input), summary… See the full description on the dataset page: https://huggingface.co/datasets/razhan/script_normalization_ckb.text1M<n<10M0 likes45 downloads1d agoHugging Face06AdamRoch /nfl-play-normalization-v2 NFL Play Normalization V2 Data-Efficiency Curve This dataset contains four nested training sets with 3,125, 6,250, 12,500, and 25,000 unchanged real nflverse play descriptions from the 2019 through 2022 seasons. The v2 selection targets laterals, repeated fumbles, accepted penalties after turnovers, and descriptions with several yardage clauses. The curve manifest pins every dataset and point-manifest checksum. The review file explains the deterministic selection and the… See the full description on the dataset page: https://huggingface.co/datasets/AdamRoch/nfl-play-normalization-v2.text-generation0 likes40 downloads29d agoHugging Face07Zarinaaa /kyrgyz-text-normalization Kyrgyz Text Normalization Dataset A dataset for training and evaluating Kyrgyz text normalization systems. Released subset accompanying "Kyrgyz Text Normalization: A Comparative Study of Neural and Rule-Based Approaches" (MeLLM Workshop @ ACL 2026). What is in this release This is a representative 20,000-pair subset of a larger 1.67M-pair training corpus, plus the full 1,000-example human-verified test set used in the paper. Split Examples Source Verification… See the full description on the dataset page: https://huggingface.co/datasets/Zarinaaa/kyrgyz-text-normalization.texttext-generation10K<n<100K0 likes39 downloads4mo agoHugging Face08mschonhardt /georges-1913-normalization Normalized Georges 1913 Description This dataset was created as part of the Burchard's Dekret Digital project (www.burchards-dekret-digital.de), funded by the Academy of Sciences and Literature | Mainz. It is based on 55,000 lemmata from Karl Georges, Ausführliches lateinisch-deutsches Handwörterbuch, Hannover 1913 (Georges 1913) and was developed to train models for normalization tasks in the context of medieval Latin. The dataset consists of approximately 5 million… See the full description on the dataset page: https://huggingface.co/datasets/mschonhardt/georges-1913-normalization.text1M<n<10M0 likes37 downloads2y agoHugging Face09sellersew /carrot-engine-normalization-translation-v2text10M<n<100M1 likes35 downloads3y agoHugging Face10yagmurtuncer /turkish-chat-normalization-mini Turkish Chat Normalization Mini turkish-chat-normalization-mini is a web-derived and rule-degraded Turkish text normalization dataset designed for rewriting noisy, informal, unpunctuated, or diacritics-missing Turkish text into cleaner and more readable Turkish. The dataset does not contain private user messages, chat logs, social media comments, complaint records, or scraped personal conversations. Source sentences are collected from open Turkish web resources, while the input… See the full description on the dataset page: https://huggingface.co/datasets/yagmurtuncer/turkish-chat-normalization-mini.texttext-generation10K<n<100K0 likes34 downloads4mo agoHugging Face11pavanBuduguppa /asr_inverse_text_normalization1 likes31 downloads4y agoHugging Face12yagmurtuncer /turkish-text-normalization 🇹🇷 Turkish Text Normalization (TN / ITN) A deterministic, rule-based dataset of Turkish written ↔ spoken pairs for Text Normalization (TN) and Inverse Text Normalization (ITN) — mapping digit/symbol forms (1.500 TL, %25, 15.07.2026) to their fully spoken Turkish words (bin beş yüz lira, yüzde yirmi beş, on beş temmuz iki bin yirmi altı) and back. This is a common, high-value preprocessing step for Turkish ASR post-processing and TTS front-ends, where numbers, dates, currencies… See the full description on the dataset page: https://huggingface.co/datasets/yagmurtuncer/turkish-text-normalization.texttext-generation10K<n<100K0 likes30 downloads2mo agoHugging Face13larrylawl /chinese-lexical-normalization chinese-lexical-normalization This dataset contains informal-formal-explanation triples from the chinese-lexical-normalization dataset. Note that there are duplicate informal-formal pairs due to multiple explanations. Example usage: from datasets import load_dataset dataset = load_dataset("larrylawl/chinese-lexical-normalization") text1K<n<10K0 likes29 downloads3y agoHugging Face14llmsql-bench /prompts_for_tables_normalization_and_new_sqlstext100K<n<1M0 likes28 downloads4mo agoHugging Face15dataautogpt3 /normalization_faces Dataset Card for "normalization_faces" More Information needed imagen<1K4 likes26 downloads3y agoHugging Face16kenenbek /gemma-russian-normalization-datasettext100K<n<1M0 likes26 downloads11mo agoHugging Face17sreetz-nv /eval_gr00t_n1d6-vials_rackleft_real_fix_normalization_0202_evalThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 2, "total_frames": 2536, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sreetz-nv/eval_gr00t_n1d6-vials_rackleft_real_fix_normalization_0202_eval.tabularrobotics1K<n<10K0 likes26 downloads8mo agoHugging Face18sreetz-nv /eval_groot-vials_rackleft_real_fix_normalization_0202_2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 0, "total_frames": 0, "total_tasks": 0, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": {}, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sreetz-nv/eval_groot-vials_rackleft_real_fix_normalization_0202_2.robotics0 likes26 downloads8mo agoHugging Face19GoktugD /turkish-text-normalization-1m Turkish Text Normalization 1M v2 Kontrollü altı gürültü türüyle Türkçe metin normalizasyon çiftleri. Doğrulanmış boyut Train: 980,000 Validation: 10,000 Test: 10,000 Toplam: 1,000,000 Ana görev sütunları: id, noisy_text, normalized_text, noise_type Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type, provenance… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-text-normalization-1m.texttext-generation1M<n<10M0 likes24 downloads1mo agoHugging Face20AdamRoch /nfl-play-normalization-2019-2022 NFL play normalization, 2019-2022 This dataset contains 188452 real raw nflverse play descriptions from 2019 through 2022. Each record pairs the raw description with a canonical JSON expected record for NFL play normalization. License and attribution The source nflverse releases are CC BY 4.0. This derived dataset retains source season, game ID, play ID, source URLs, and SHA-256 checksums in metadata/training-manifest.json. Attribute nflverse when using this… See the full description on the dataset page: https://huggingface.co/datasets/AdamRoch/nfl-play-normalization-2019-2022.text-generation0 likes24 downloads1mo agoHugging Face21sreetz-nv /eval_groot-vials_rackleft_real_fix_normalization_0202This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 0, "total_frames": 0, "total_tasks": 0, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": {}, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sreetz-nv/eval_groot-vials_rackleft_real_fix_normalization_0202.robotics0 likes23 downloads8mo agoHugging Face22jaio98 /basque_dialect_normalizationtext1K<n<10K0 likes23 downloads4mo agoHugging Face23sreetz-nv /eval_groot-vials_rackleft_real_fix_normalization_0202_8This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 1, "total_frames": 506, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sreetz-nv/eval_groot-vials_rackleft_real_fix_normalization_0202_8.tabularroboticsn<1K0 likes22 downloads8mo agoHugging Face24LegionIntel /date_string_normalizationtext10K<n<100K0 likes20 downloads2y agoHugging Face25distilabel-internal-testing /deita-no-normalization Dataset Card for deita-no-normalization This dataset has been created with Distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/distilabel-internal-testing/deita-no-normalization/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/distilabel-internal-testing/deita-no-normalization.tabular1K<n<10K0 likes19 downloads2y agoHugging Face26Yasshhhh /adaption-telugu-normalization This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-telugu_normalization This dataset contains a collection of user queries and statements written in Telugu, covering diverse topics such as mobile troubleshooting, cyber security threats, insurance renewals, and emergency services. The samples vary from short keywords to detailed problem descriptions involving scams, natural disasters, and administrative procedures. It represents… See the full description on the dataset page: https://huggingface.co/datasets/Yasshhhh/adaption-telugu-normalization.text1K<n<10K0 likes19 downloads3mo agoHugging Face27psidharth567 /normalizationtext100K<n<1M0 likes18 downloads11mo agoHugging Face28DigitalUmuganda /Text_Normalization_Challenge_Unittests_Eng_Fratextn<1K0 likes17 downloads11mo agoHugging Face29GoktugD /turkish-datetime-normalization-500k Turkish Datetime Normalization 500K v2 Türkçe tarih-saat ifadelerini ISO-8601 ve Europe/Istanbul saat dilimine eşler. Doğrulanmış boyut Train: 490,000 Validation: 5,000 Test: 5,000 Toplam: 500,000 Ana görev sütunları: id, text, normalized_datetime, timezone Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type, provenance… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-datetime-normalization-500k.texttext-generation100K<n<1M0 likes17 downloads1mo agoHugging Face30huseinzol05 /filtered-common-crawl-abstractive-normalization0 likes14 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.