lexic
Datasets
All datasets matching “lexic”arabic-stem-lexicon
Arabic Diacritized-Stem Lexicon
An undiacritized Arabic surface form → its most frequent diacritized stem.
Standard Arabic writes no short vowels, so anything that has to pronounce Arabic
must first put them back. A neural diacritizer does that well on rare words, where
inference is the only thing there is. On common words it is the wrong tool:
which vowels كتاب carries is not a thing to be inferred, it is a thing to be looked
up — and models get exactly these wrong, reading… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-stem-lexicon.MATH_qCoT_LLMquery_questionasquery_lexicalqueryDatasets from Paper: https://huggingface.co/papers/2505.18405
portuguese-unified-pronunciation-lexicon
Portuguese Unified Pronunciation Lexicon
A flat, single-row-per-pronunciation dataset merging Portuguese IPA transcriptions from three authoritative sources. Each row is a word × region × POS tuple with both broad phonemic (ipa_broad) and narrow phonetic (ipa_narrow) transcriptions normalized across sources.
Source
Words
Convention
Description
Infopédia (Porto Editora)
102,685
Broad phonemic
European Portuguese dictionary IPA
Wiktionary (pt.wiktionary.org)
15,720… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-unified-pronunciation-lexicon.jastrow-lexicon
NuBerea/jastrow-lexicon
A complete Public-Domain ingest of Marcus Jastrow's A Dictionary of the Targumim,
the Talmud Babli and Yerushalmi, and the Midrashic Literature (London/New York,
1903) — the standard reference lexicon for the language of the Tannaitic, Amoraic,
Targumic, and Midrashic corpora. It is the canonical PD witness lexicon for
Tannaitic Hebrew, Targumic Aramaic (Onqelos / Yerushalmi / Pseudo-Jonathan),
Babylonian Talmudic Aramaic, and Palestinian/Galilean Aramaic… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/jastrow-lexicon.distributional-lexicon
Distributional Sense Lexicon — Greek (archaic → Byzantine) + Latin + Hebrew
A corpus-driven sense lexicon of ancient Greek, Latin, and Biblical Hebrew that represents
word meaning directly as the distribution of attested usage — graded sense membership,
frequency mass, collocational profile, and diachronic drift — scaffolded on the Louw–Nida
sense taxonomy. The Greek arm spans the full diachronic range from archaic/classical
literature through the New Testament, Septuagint… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/distributional-lexicon.multilingual_toxic_lexicon
Multilingual Toxic Lexicon
[2025] The lexicon is extended to new languages! Now also included: Italian, French, Hebrew, Hindi, Japanese, Tatar. The list is used on TextDetox 2025 shared task.
[2024] The compilation for 9 languages (English, Russian, Ukrainian, Spanish, German, Amharic, Arabic, Chinese, Hindi) toxic words lists which is used for TextDetox 2024 shared task.
The list of original sources:
English: link
Russian: link
Ukrainian: link
Spanish: link
German: link
Amhairc:… See the full description on the dataset page: https://huggingface.co/datasets/textdetox/multilingual_toxic_lexicon.
