normalization
yoruba-normalization-pairs
Normalization pairs dataset
What this is
24,475 pairs of Yorùbá text, each a corrupted form next to its canonical form, labelled by corruption type. I built it for testing orthographic normalization code.
The library
This dataset was built alongside yotext, a Python library for Yorùbá orthographic normalization and diacritic restoration. The library is on PyPI at https://pypi.org/project/yotext/ and the source is at… See the full description on the dataset page: https://huggingface.co/datasets/adedejimakinde/yoruba-normalization-pairs.uk-text-normalization
Український TTS-нормалізатор — датасет
Пари «письмовий текст → як його вимовляють» для української. Числа, дати, час,
гроші, одиниці, скорочення, коди, телефони, IBAN, домени, пошта, римські
цифри, латинські вкраплення — те, що треба розгорнути словами перед синтезом
мовлення.
{"task_id": 0,
"combo_names": ["Кількісні числівники (написані цифрами)",
"Порядкові числівники (написані цифрами з закінченням)"],
"original": "На 1-й полиці стоять 4 книги."… See the full description on the dataset page: https://huggingface.co/datasets/skypro1111/uk-text-normalization.bangla-dialect-normalization
Bangla Dialect Normalization Dataset
A parallel corpus mapping standard Bangla to five regional Bangla dialects,
built from the Vashantor dataset. Each row contains the same sentence in
standard Bangla and Banglish (romanized), alongside its dialect Bangla and
dialect Banglish equivalent, plus an English gloss.
Regions covered
Barishal, Chittagong, Mymensingh, Noakhali, Sylhet
Schema
Field
Description
standard_bangla
Sentence in standard… See the full description on the dataset page: https://huggingface.co/datasets/zmsali/bangla-dialect-normalization.config-normalization-0909Controlled path-normalization probe. No third-party data.
script_normalization_ckb
Script normalization CKB — noisy → standard Sorani (script_normalization_ckb)
Nearly 6M sentence pairs: the text column holds Central Kurdish written with
non-standard or distorted characters, summary holds the same sentence in standard
Sorani orthography. Use it for text normalisation, spell correction, or to learn the
character-level mapping rules.
At a glance
Rows
5,997,025 — train 5,330,689 / test 666,336
Columns
text (noisy input), summary… See the full description on the dataset page: https://huggingface.co/datasets/razhan/script_normalization_ckb.nfl-play-normalization-v2
NFL Play Normalization V2 Data-Efficiency Curve
This dataset contains four nested training sets with 3,125, 6,250, 12,500,
and 25,000 unchanged real nflverse play descriptions from the 2019 through
2022 seasons. The v2 selection targets laterals, repeated fumbles, accepted
penalties after turnovers, and descriptions with several yardage clauses.
The curve manifest pins every dataset and point-manifest checksum. The review
file explains the deterministic selection and the… See the full description on the dataset page: https://huggingface.co/datasets/AdamRoch/nfl-play-normalization-v2.
