CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mkd-chanwoo /normalized-datasets-for-koreanLLM Normalized Datasets for Korean LLM (Stage 0.5) This is a comprehensive pretraining dataset for Korean LLMs containing ~944M normalized documents from 42 source datasets across English, Korean, Code, and Science domains. Quick Stats Metric Value Total Documents ~944,545,628 Source Datasets 42 Format JSONL (one JSON object per line) Languages English, Korean Domains English, Korean, Code, Science Pipeline Stage Stage 0.5 (after download, before… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/normalized-datasets-for-koreanLLM.text1M<n<10M1 likes3.7k downloads4mo agoHugging Face02Scicom-intl /Multilingual-Normalizer Multilingual TTS text normalizer (written → spoken) Training pairs for fine-tuning a small LLM as a text-to-speech normalizer: text is a sentence the way people type it (digits, currency symbols, dates, phone numbers, …) and normalized is the exact spoken form, in the same language, with nothing left that a TTS model cannot say. 52,698 rows — 16 monolingual locales and 6 Malaysian code-switched pairs. Every row is digit-free on the spoken side. from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Multilingual-Normalizer.texttext-generation100K<n<1M0 likes203 downloads16d agoHugging Face03adedejimakinde /yoruba-normalization-pairs Normalization pairs dataset What this is 24,475 pairs of Yorùbá text, each a corrupted form next to its canonical form, labelled by corruption type. I built it for testing orthographic normalization code. The library This dataset was built alongside yotext, a Python library for Yorùbá orthographic normalization and diacritic restoration. The library is on PyPI at https://pypi.org/project/yotext/ and the source is at… See the full description on the dataset page: https://huggingface.co/datasets/adedejimakinde/yoruba-normalization-pairs.texttext-generation10K<n<100K1 likes100 downloads11d agoHugging Face04skypro1111 /uk-text-normalization Український TTS-нормалізатор — датасет Пари «письмовий текст → як його вимовляють» для української. Числа, дати, час, гроші, одиниці, скорочення, коди, телефони, IBAN, домени, пошта, римські цифри, латинські вкраплення — те, що треба розгорнути словами перед синтезом мовлення. {"task_id": 0, "combo_names": ["Кількісні числівники (написані цифрами)", "Порядкові числівники (написані цифрами з закінченням)"], "original": "На 1-й полиці стоять 4 книги."… See the full description on the dataset page: https://huggingface.co/datasets/skypro1111/uk-text-normalization.texttext-generation1K<n<10K2 likes77 downloads1mo agoHugging Face05francescortu /DistillDetect-normalized-traces DistillDetect — format-normalized teacher traces Teacher responses from Reference-Based Distillation Detection in LLMs (arXiv:2607.09692), rewritten so that every teacher uses the same output format. 7,918 rows across 8 teacher/prompt-set pairs. Why this exists In the released data each teacher emits a structurally different response, so a student trained on it — and any detector trained to attribute it — can key on surface format instead of the teacher's actual… See the full description on the dataset page: https://huggingface.co/datasets/francescortu/DistillDetect-normalized-traces.texttext-generation1K<n<10K0 likes77 downloads28d agoHugging Face06zmsali /bangla-dialect-normalization Bangla Dialect Normalization Dataset A parallel corpus mapping standard Bangla to five regional Bangla dialects, built from the Vashantor dataset. Each row contains the same sentence in standard Bangla and Banglish (romanized), alongside its dialect Bangla and dialect Banglish equivalent, plus an English gloss. Regions covered Barishal, Chittagong, Mymensingh, Noakhali, Sylhet Schema Field Description standard_bangla Sentence in standard… See the full description on the dataset page: https://huggingface.co/datasets/zmsali/bangla-dialect-normalization.texttranslation10K<n<100K0 likes65 downloads22d agoHugging Face07normalcomputing /wikiqa-counterfactualModel Card for Long-range Counterfactual WikiQA Github: https://github.com/normal-computing/extended-mind-transformers/ ArXiv: https://arxiv.org/abs/2406.02332 Original dataset by Abacus AI. Developed by: Normal Computing, Adapted from Abacus AI License: Apache 2.0 Long-range Counterfactual Retrieval Benchmark This benchmark is a modified wikiQA benchmark. The dataset is composed of Wikipedia articles (of 2-16 thousand tokens) and corresponding questions. We modify the… See the full description on the dataset page: https://huggingface.co/datasets/normalcomputing/wikiqa-counterfactual.textn<1K1 likes50 downloads2y agoHugging Face08lemon-mint /OpenThoughts-114k-Normalizedprefixes = [ "Generate an executable Python function generated from the given prompt. The function should take stdin as input and print the output. Simply call the function after the definition.", "Return your final response within \\boxed{}. ", "Generate an executable Python function generated from the given prompt. Return the function body without invoking it at the final solution.", ] // -1 if None text100K<n<1M1 likes49 downloads2y agoHugging Face09atrevidasadia /dair-ai-emotion-normalized-instruction-input-output dair-ai emotion | normalized Summary Dataset ID: 143 Type: normalized Rows: 16,000 Source: dair-ai/emotion Dataset Sources #143 dair-ai emotion | normalized [normalized | 16,000 rows] Notes Edited and Exported from the Kitsune Training Suite (Forge) Review the dataset artifact and metadata before publishing. Citation > via dair-ai @inproceedings{saravia-etal-2018-carer, title = "{CARER}: Contextualized Affect… See the full description on the dataset page: https://huggingface.co/datasets/atrevidasadia/dair-ai-emotion-normalized-instruction-input-output.texttext-generation10K<n<100K0 likes44 downloads28d agoHugging Face10Zarinaaa /kyrgyz-text-normalization Kyrgyz Text Normalization Dataset A dataset for training and evaluating Kyrgyz text normalization systems. Released subset accompanying "Kyrgyz Text Normalization: A Comparative Study of Neural and Rule-Based Approaches" (MeLLM Workshop @ ACL 2026). What is in this release This is a representative 20,000-pair subset of a larger 1.67M-pair training corpus, plus the full 1,000-example human-verified test set used in the paper. Split Examples Source Verification… See the full description on the dataset page: https://huggingface.co/datasets/Zarinaaa/kyrgyz-text-normalization.texttext-generation10K<n<100K0 likes39 downloads4mo agoHugging Face11cometadata /2025-08-datacite-normalized-affiliation-string-distribution DataCite Normalized Affiliation Distribution Summary normalized_distribution.json contains one JSON object per normalized affiliation string. It aggregates the total occurrence count, a ranked list of the raw affiliation strings that collapse into the normalized form, and the provider/client entities that asserted them. This dataset is derived from the August 2025 DataCite creator/contributor export. Structure { "normalized": "example university"… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/2025-08-datacite-normalized-affiliation-string-distribution.text1M<n<10M0 likes30 downloads11mo agoHugging Face12llmsql-bench /prompts_for_tables_normalization_and_new_sqlstext100K<n<1M0 likes28 downloads4mo agoHugging Face13deltakitsune /dair-ai-emotion-normalized-instruction-input-output dair-ai emotion | normalized Summary Dataset ID: 143 Type: normalized Rows: 16,000 Source: dair-ai/emotion Dataset Sources #143 dair-ai emotion | normalized [normalized | 16,000 rows] Notes Edited and Exported from the Kitsune Training Suite (Forge) Review the dataset artifact and metadata before publishing. Citation > via dair-ai @inproceedings{saravia-etal-2018-carer, title = "{CARER}: Contextualized Affect Representations for… See the full description on the dataset page: https://huggingface.co/datasets/deltakitsune/dair-ai-emotion-normalized-instruction-input-output.texttext-generation10K<n<100K0 likes27 downloads5mo agoHugging Face14llm-semantic-router /halueval-spans-normalized HaluEval Span-Level Dataset (RAGTruth-Normalized Prompts) 🔍 Span-level hallucination detection dataset with prompts normalized to match RAGTruth format for improved cross-dataset compatibility. Quick Start from datasets import load_dataset dataset = load_dataset("llm-semantic-router/halueval-spans-normalized") Why Normalized Prompts? Training on mixed datasets with different prompt formats causes distribution shift: Original Format Normalized Format… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/halueval-spans-normalized.texttoken-classification10K<n<100K0 likes26 downloads9mo agoHugging Face15nassimjp /normal_stories_translated_in_pashtotextn<1K0 likes20 downloads1mo agoHugging Face16Yasshhhh /adaption-telugu-normalization This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-telugu_normalization This dataset contains a collection of user queries and statements written in Telugu, covering diverse topics such as mobile troubleshooting, cyber security threats, insurance renewals, and emergency services. The samples vary from short keywords to detailed problem descriptions involving scams, natural disasters, and administrative procedures. It represents… See the full description on the dataset page: https://huggingface.co/datasets/Yasshhhh/adaption-telugu-normalization.text1K<n<10K0 likes19 downloads3mo agoHugging Face17PhdDz /PubmedQA_5_WITH_RELATION_vsimilarity_primekg_normalizedtext1K<n<10K0 likes18 downloads2y agoHugging Face18Kicshikxo /dataset-context-self.qwen2.5-max2048.normalizedtext1M<n<10M0 likes18 downloads2mo agoHugging Face19PhdDz /PubmedQA_5_WITH_RELATION_vsimilarity_both_normalizedtext1K<n<10K0 likes15 downloads2y agoHugging Face20PhdDz /PubmedQA_5_WITH_RELATION_vqc_primekg_normalizedtext1K<n<10K0 likes14 downloads2y agoHugging Face21Jnx03 /kanitakorn-deepseek-v44-normalized-mcq-replay-mix Kanitakorn v44 normalized MCQ replay mix Original and previously audited synthetic SFT mixture for a non-Thai-base DeepSeek/Qwen-style <=14B candidate. This dataset does not include benchmark prompts, benchmark gold answers, model benchmark samples, BoN traces, routing labels, or consensus outputs. Design intent: normalize MCQ final marker to คำตอบคือ (x) keep explanations before the answer instead of answer-only rows target aggregate ThaiExam failure buckets: grammar/spelling… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v44-normalized-mcq-replay-mix.text1K<n<10K0 likes13 downloads3mo agoHugging Face22backup-dev /normalize_symlinkstabularn<1K0 likes12 downloads5mo agoHugging Face23ozgurkrkrt /zemberek-normalizetext100K<n<1M0 likes10 downloads2y agoHugging Face24lunahr /normalization-data-mixed Normalization Dataset (Mixed) This dataset is a collection of 50000 rows originating from various sources: Wikipedia - 20000 rows PersonaChat truecased - 20000 rows Synthetic edge case data - 5000 rows Synthetic quoted text data - 5000 rows The synthetic data has been generated using GPT-5.3 models. The other data was sourced from the original Hugging Face sources. This dataset can be used to train text normalizers that convert badly formatted English into correct English.… See the full description on the dataset page: https://huggingface.co/datasets/lunahr/normalization-data-mixed.text10K<n<100K0 likes10 downloads2mo agoHugging Face25brendan-gho /qwen3b_normal_numstext10K<n<100K0 likes10 downloads5mo agoHugging Face26joduor /adaption-gd-unk-normalized-samples This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-gd_unk_normalized_samples This dataset contains normalized and semantically enriched samples identified by GD-UNK codes, featuring anomaly labels and timestamps. The content is structured to ensure consistency and readiness for machine learning models, avoiding hallucinations. Each entry includes object data points processed for quality enhancement. Dataset size There are 1… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-gd-unk-normalized-samples.textn<1K0 likes10 downloads5mo agoHugging Face27happy8825 /after_incident_normaltextn<1K0 likes9 downloads9mo agoHugging Face28PhdDz /PubmedQA_5_WITH_DEFINITION_RELATION_vsimilarity_primekg_normalizedtext1K<n<10K0 likes8 downloads2y agoHugging Face29open-llm-leaderboard /ehristoforu__fq2.5-7b-it-normalize_false-detailsgated Dataset Card for Evaluation run of ehristoforu/fq2.5-7b-it-normalize_false Dataset automatically created during the evaluation run of model ehristoforu/fq2.5-7b-it-normalize_false The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ehristoforu__fq2.5-7b-it-normalize_false-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face30ruscorpora /normalization TL;DR: Text Normalization for Social Media Corpus Dataset Description This dataset contains examples of Russian-language texts from social networks with distorted spelling (typos, abbreviations, etc.) and their normalized versions in json format. A detailed spelling correction protocol is given in the TBA article. The dataset size is 1930 sentence pairs. In each pair, the sentences are tokenized by words, and the lengths of both sentences in the pair are equal. If a… See the full description on the dataset page: https://huggingface.co/datasets/ruscorpora/normalization.text1K<n<10K0 likes8 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.