datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
normalized-datasets-for-koreanLLM
Normalized Datasets for Korean LLM (Stage 0.5)
This is a comprehensive pretraining dataset for Korean LLMs containing ~944M normalized documents from 42 source datasets across English, Korean, Code, and Science domains.
Quick Stats
Metric
Value
Total Documents
~944,545,628
Source Datasets
42
Format
JSONL (one JSON object per line)
Languages
English, Korean
Domains
English, Korean, Code, Science
Pipeline Stage
Stage 0.5 (after download, before… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/normalized-datasets-for-koreanLLM.Multilingual-Normalizer
Multilingual TTS text normalizer (written → spoken)
Training pairs for fine-tuning a small LLM as a text-to-speech normalizer: text is a sentence
the way people type it (digits, currency symbols, dates, phone numbers, …) and normalized is the
exact spoken form, in the same language, with nothing left that a TTS model cannot say.
52,698 rows — 16 monolingual locales and 6 Malaysian code-switched pairs. Every row is
digit-free on the spoken side.
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Multilingual-Normalizer.yoruba-normalization-pairs
Normalization pairs dataset
What this is
24,475 pairs of Yorùbá text, each a corrupted form next to its canonical form, labelled by corruption type. I built it for testing orthographic normalization code.
The library
This dataset was built alongside yotext, a Python library for Yorùbá orthographic normalization and diacritic restoration. The library is on PyPI at https://pypi.org/project/yotext/ and the source is at… See the full description on the dataset page: https://huggingface.co/datasets/adedejimakinde/yoruba-normalization-pairs.uk-text-normalization
Український TTS-нормалізатор — датасет
Пари «письмовий текст → як його вимовляють» для української. Числа, дати, час,
гроші, одиниці, скорочення, коди, телефони, IBAN, домени, пошта, римські
цифри, латинські вкраплення — те, що треба розгорнути словами перед синтезом
мовлення.
{"task_id": 0,
"combo_names": ["Кількісні числівники (написані цифрами)",
"Порядкові числівники (написані цифрами з закінченням)"],
"original": "На 1-й полиці стоять 4 книги."… See the full description on the dataset page: https://huggingface.co/datasets/skypro1111/uk-text-normalization.DistillDetect-normalized-traces
DistillDetect — format-normalized teacher traces
Teacher responses from Reference-Based Distillation Detection in LLMs
(arXiv:2607.09692), rewritten so that every
teacher uses the same output format. 7,918 rows across 8 teacher/prompt-set
pairs.
Why this exists
In the released data each teacher emits a structurally different response, so a
student trained on it — and any detector trained to attribute it — can key on
surface format instead of the teacher's actual… See the full description on the dataset page: https://huggingface.co/datasets/francescortu/DistillDetect-normalized-traces.bangla-dialect-normalization
Bangla Dialect Normalization Dataset
A parallel corpus mapping standard Bangla to five regional Bangla dialects,
built from the Vashantor dataset. Each row contains the same sentence in
standard Bangla and Banglish (romanized), alongside its dialect Bangla and
dialect Banglish equivalent, plus an English gloss.
Regions covered
Barishal, Chittagong, Mymensingh, Noakhali, Sylhet
Schema
Field
Description
standard_bangla
Sentence in standard… See the full description on the dataset page: https://huggingface.co/datasets/zmsali/bangla-dialect-normalization.wikiqa-counterfactualModel Card for Long-range Counterfactual WikiQA
Github: https://github.com/normal-computing/extended-mind-transformers/
ArXiv: https://arxiv.org/abs/2406.02332
Original dataset by Abacus AI.
Developed by: Normal Computing, Adapted from Abacus AI
License: Apache 2.0
Long-range Counterfactual Retrieval Benchmark
This benchmark is a modified wikiQA benchmark. The dataset is composed of Wikipedia articles (of 2-16 thousand tokens) and corresponding questions. We modify the… See the full description on the dataset page: https://huggingface.co/datasets/normalcomputing/wikiqa-counterfactual.OpenThoughts-114k-Normalizedprefixes = [
"Generate an executable Python function generated from the given prompt. The function should take stdin as input and print the output. Simply call the function after the definition.",
"Return your final response within \\boxed{}. ",
"Generate an executable Python function generated from the given prompt. Return the function body without invoking it at the final solution.",
]
// -1 if None
dair-ai-emotion-normalized-instruction-input-output
dair-ai emotion | normalized
Summary
Dataset ID: 143
Type: normalized
Rows: 16,000
Source: dair-ai/emotion
Dataset Sources
#143 dair-ai emotion | normalized [normalized | 16,000 rows]
Notes
Edited and Exported from the Kitsune Training Suite (Forge)
Review the dataset artifact and metadata before publishing.
Citation > via dair-ai
@inproceedings{saravia-etal-2018-carer,
title = "{CARER}: Contextualized Affect… See the full description on the dataset page: https://huggingface.co/datasets/atrevidasadia/dair-ai-emotion-normalized-instruction-input-output.kyrgyz-text-normalization
Kyrgyz Text Normalization Dataset
A dataset for training and evaluating Kyrgyz text normalization systems. Released subset accompanying "Kyrgyz Text Normalization: A Comparative Study of Neural and Rule-Based Approaches" (MeLLM Workshop @ ACL 2026).
What is in this release
This is a representative 20,000-pair subset of a larger 1.67M-pair training corpus, plus the full 1,000-example human-verified test set used in the paper.
Split
Examples
Source
Verification… See the full description on the dataset page: https://huggingface.co/datasets/Zarinaaa/kyrgyz-text-normalization.2025-08-datacite-normalized-affiliation-string-distribution
DataCite Normalized Affiliation Distribution
Summary
normalized_distribution.json contains one JSON object per normalized affiliation string. It aggregates the total
occurrence count, a ranked list of the raw affiliation strings that collapse into the normalized form, and the
provider/client entities that asserted them. This dataset is derived from the August 2025 DataCite creator/contributor export.
Structure
{
"normalized": "example university"… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/2025-08-datacite-normalized-affiliation-string-distribution.prompts_for_tables_normalization_and_new_sqlsdair-ai-emotion-normalized-instruction-input-output
dair-ai emotion | normalized
Summary
Dataset ID: 143
Type: normalized
Rows: 16,000
Source: dair-ai/emotion
Dataset Sources
#143 dair-ai emotion | normalized [normalized | 16,000 rows]
Notes
Edited and Exported from the Kitsune Training Suite (Forge)
Review the dataset artifact and metadata before publishing.
Citation > via dair-ai
@inproceedings{saravia-etal-2018-carer,
title = "{CARER}: Contextualized Affect Representations for… See the full description on the dataset page: https://huggingface.co/datasets/deltakitsune/dair-ai-emotion-normalized-instruction-input-output.halueval-spans-normalized
HaluEval Span-Level Dataset (RAGTruth-Normalized Prompts)
🔍 Span-level hallucination detection dataset with prompts normalized to match RAGTruth format for improved cross-dataset compatibility.
Quick Start
from datasets import load_dataset
dataset = load_dataset("llm-semantic-router/halueval-spans-normalized")
Why Normalized Prompts?
Training on mixed datasets with different prompt formats causes distribution shift:
Original Format
Normalized Format… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/halueval-spans-normalized.normal_stories_translated_in_pashtoadaption-telugu-normalization
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-telugu_normalization
This dataset contains a collection of user queries and statements written in Telugu, covering diverse topics such as mobile troubleshooting, cyber security threats, insurance renewals, and emergency services. The samples vary from short keywords to detailed problem descriptions involving scams, natural disasters, and administrative procedures. It represents… See the full description on the dataset page: https://huggingface.co/datasets/Yasshhhh/adaption-telugu-normalization.PubmedQA_5_WITH_RELATION_vsimilarity_primekg_normalizeddataset-context-self.qwen2.5-max2048.normalizedPubmedQA_5_WITH_RELATION_vsimilarity_both_normalizedPubmedQA_5_WITH_RELATION_vqc_primekg_normalizedkanitakorn-deepseek-v44-normalized-mcq-replay-mix
Kanitakorn v44 normalized MCQ replay mix
Original and previously audited synthetic SFT mixture for a non-Thai-base DeepSeek/Qwen-style <=14B candidate.
This dataset does not include benchmark prompts, benchmark gold answers, model benchmark samples, BoN traces, routing labels, or consensus outputs.
Design intent:
normalize MCQ final marker to คำตอบคือ (x)
keep explanations before the answer instead of answer-only rows
target aggregate ThaiExam failure buckets: grammar/spelling… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v44-normalized-mcq-replay-mix.normalize_symlinkszemberek-normalizenormalization-data-mixed
Normalization Dataset (Mixed)
This dataset is a collection of 50000 rows originating from various sources:
Wikipedia - 20000 rows
PersonaChat truecased - 20000 rows
Synthetic edge case data - 5000 rows
Synthetic quoted text data - 5000 rows
The synthetic data has been generated using GPT-5.3 models.
The other data was sourced from the original Hugging Face sources.
This dataset can be used to train text normalizers that convert badly formatted English into correct English.… See the full description on the dataset page: https://huggingface.co/datasets/lunahr/normalization-data-mixed.qwen3b_normal_numsadaption-gd-unk-normalized-samples
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-gd_unk_normalized_samples
This dataset contains normalized and semantically enriched samples identified by GD-UNK codes, featuring anomaly labels and timestamps. The content is structured to ensure consistency and readiness for machine learning models, avoiding hallucinations. Each entry includes object data points processed for quality enhancement.
Dataset size
There are 1… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-gd-unk-normalized-samples.after_incident_normalPubmedQA_5_WITH_DEFINITION_RELATION_vsimilarity_primekg_normalizedehristoforu__fq2.5-7b-it-normalize_false-details
Dataset Card for Evaluation run of ehristoforu/fq2.5-7b-it-normalize_false
Dataset automatically created during the evaluation run of model ehristoforu/fq2.5-7b-it-normalize_false
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ehristoforu__fq2.5-7b-it-normalize_false-details.normalization
TL;DR: Text Normalization for Social Media Corpus
Dataset Description
This dataset contains examples of Russian-language texts from social networks
with distorted spelling (typos, abbreviations, etc.) and their normalized versions
in json format. A detailed spelling correction protocol is given in the TBA article.
The dataset size is 1930 sentence pairs. In each pair, the sentences are tokenized
by words, and the lengths of both sentences in the pair are equal. If a… See the full description on the dataset page: https://huggingface.co/datasets/ruscorpora/normalization.
