CoolFace
Datasetpublic

Mikhailo/ubertext-itn-uk

Ukrainian ITN Dataset ~889k sentence pairs for Ukrainian Inverse Text Normalization (ITN). Spoken-form → Written-form pairs generated from skypro1111/ubertext-2-news-verbalized using LLM-assisted annotation. Columns spoken: verbalized Ukrainian text (input for ITN) written: normalized text with digits, symbols, abbreviations Example spoken written сорок два відсотки населення 42% населення пʼятнадцять доларів і двадцять центів $15.20… See the full description on the dataset page: https://huggingface.co/datasets/Mikhailo/ubertext-itn-uk.

sourceHugging Facecc-by-4.0updated 3mo agoView on Hugging Face
0likes7downloads

Mikhailo/ubertext-itn-uk · main · files are served by the source, never re-hosted here