datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bifonia-pt-homographs
bifonia — Portuguese Heterophonic Homograph Disambiguation
Labelled European-Portuguese (pt-PT) sentences for 27 heterophonic homographs —
words with identical spelling whose pronunciation (IPA) depends on part of speech
or meaning, e.g. para (preposition ˈpɐɾɐ vs verb ˈpaɾɐ), molho
(sauce ˈmoʎu vs bundle ˈmɔʎu), corte (royal court ˈkoɾtɨ vs cut ˈkɔɾtɨ).
Useful for grapheme-to-phoneme / TTS front-ends and for POS disambiguation.
Schema
field
description… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/bifonia-pt-homographs.HomographResolutionEval
Homograph Resolution Evaluation Dataset
This dataset is designed to evaluate the performance of Text-to-Speech (TTS) systems and Language Models (LLMs) in resolving homographs in the Russian language. It contains carefully curated sentences, each featuring at least one homograph with the correct stress indicated. The dataset is particularly useful for assessing stress assignment tasks in TTS systems and LLMs.
Key Features
Language: Russian
Focus: Homograph resolution and… See the full description on the dataset page: https://huggingface.co/datasets/inkoziev/HomographResolutionEval.bulgarian_homographsБългарски Омографи с IPA и Ударения (1016 думи)
(Ако откриете грешки моля пишете в Community. Благодаря! Щом подготвя ще има и нови)
(Има и омографи в неформатиран вид 9800 думи в който са включени и обработените.
Няма цензурирани, но токсичните няма да бъдат описани макар да са част от БАН и БЕРОН)
За повече информация виж - Zakonite_na_Omografa.md
Описание
Този набор от данни съдържа 1016 уникални български омографа – думи, които се изписват еднакво, но се произнасят различно поради… See the full description on the dataset page: https://huggingface.co/datasets/batvanio12/bulgarian_homographs.bifonia-pt-homographs-gold
bifonia — European-Portuguese Heterophonic Homographs (Gold)
A sense-balanced, independently-labelled real-text benchmark for
disambiguating European-Portuguese heterophonic homographs: words spelled
identically whose pronunciation depends on meaning, not just part of speech.
sede is thirst (ˈsedɨ, closed e) or headquarters (ˈsɛdɨ, open e);
forma is a mould (ˈfoɾmɐ) or a shape (ˈfɔɾmɐ); molho is sauce
(ˈmoʎu) or a bundle (ˈmɔʎu). A text-to-speech front-end that guesses wrong… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/bifonia-pt-homographs-gold.homograph-bel-contextual-v1
HomographBel Contextual v1
Prepared contexts for fine-tuning and evaluating contextual Belarusian homograph stress
resolution. Each row describes one exact target occurrence and preserves its source text,
half-open character span, dictionary identifiers, provenance, quality tier, orthography,
grouped split, and sampling weight.
This is the training dataset for
fosters/homograph-bel-xlm-roberta-base.
Dataset summary
Dataset ID:… See the full description on the dataset page: https://huggingface.co/datasets/fosters/homograph-bel-contextual-v1.bifonia-pt-homographs-wild
bifonia-pt-homographs-wild — real-world OOD evaluation set
Real Portuguese sentences containing 112 heterophonic homographs — words spelled
identically whose pronunciation (IPA) depends on meaning, e.g. sede
(thirst ˈsedɨ vs seat ˈsɛdɨ, both nouns), forma (mould ˈfoɾmɐ vs shape
ˈfɔɾmɐ), molho (sauce ˈmoʎu vs bundle ˈmɔʎu).
Unlike the synthetic TigreGotico/bifonia-pt-homographs
train/test set (LLM-generated), every sentence here is mined from real-world text.
It is meant as an… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/bifonia-pt-homographs-wild.
