CoolFace
20 results

lexic

TigreGotico /arabic-stem-lexicon Arabic Diacritized-Stem Lexicon An undiacritized Arabic surface form → its most frequent diacritized stem. Standard Arabic writes no short vowels, so anything that has to pronounce Arabic must first put them back. A neural diacritizer does that well on rare words, where inference is the only thing there is. On common words it is the wrong tool: which vowels كتاب carries is not a thing to be inferred, it is a thing to be looked up — and models get exactly these wrong, reading… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-stem-lexicon.tabulartext-to-speech100K<n<1M0 likes3k downloads2mo agoHugging FaceRaderspace /MATH_qCoT_LLMquery_questionasquery_lexicalqueryDatasets from Paper: https://huggingface.co/papers/2505.18405 text10K<n<100K2 likes679 downloads1y agoHugging FaceTigreGotico /portuguese-unified-pronunciation-lexicon Portuguese Unified Pronunciation Lexicon A flat, single-row-per-pronunciation dataset merging Portuguese IPA transcriptions from three authoritative sources. Each row is a word × region × POS tuple with both broad phonemic (ipa_broad) and narrow phonetic (ipa_narrow) transcriptions normalized across sources. Source Words Convention Description Infopédia (Porto Editora) 102,685 Broad phonemic European Portuguese dictionary IPA Wiktionary (pt.wiktionary.org) 15,720… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-unified-pronunciation-lexicon.texttext-generation100K<n<1M1 likes540 downloads2mo agoHugging FaceNuBerea /jastrow-lexicongated NuBerea/jastrow-lexicon A complete Public-Domain ingest of Marcus Jastrow's A Dictionary of the Targumim, the Talmud Babli and Yerushalmi, and the Midrashic Literature (London/New York, 1903) — the standard reference lexicon for the language of the Tannaitic, Amoraic, Targumic, and Midrashic corpora. It is the canonical PD witness lexicon for Tannaitic Hebrew, Targumic Aramaic (Onqelos / Yerushalmi / Pseudo-Jonathan), Babylonian Talmudic Aramaic, and Palestinian/Galilean Aramaic… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/jastrow-lexicon.tabular10K<n<100K0 likes433 downloads2mo agoHugging FaceNuBerea /distributional-lexicongated Distributional Sense Lexicon — Greek (archaic → Byzantine) + Latin + Hebrew A corpus-driven sense lexicon of ancient Greek, Latin, and Biblical Hebrew that represents word meaning directly as the distribution of attested usage — graded sense membership, frequency mass, collocational profile, and diachronic drift — scaffolded on the Louw–Nida sense taxonomy. The Greek arm spans the full diachronic range from archaic/classical literature through the New Testament, Septuagint… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/distributional-lexicon.tabular10M<n<100M0 likes394 downloads4d agoHugging Facetextdetox /multilingual_toxic_lexicon Multilingual Toxic Lexicon [2025] The lexicon is extended to new languages! Now also included: Italian, French, Hebrew, Hindi, Japanese, Tatar. The list is used on TextDetox 2025 shared task. [2024] The compilation for 9 languages (English, Russian, Ukrainian, Spanish, German, Amharic, Arabic, Chinese, Hindi) toxic words lists which is used for TextDetox 2024 shared task. The list of original sources: English: link Russian: link Ukrainian: link Spanish: link German: link Amhairc:… See the full description on the dataset page: https://huggingface.co/datasets/textdetox/multilingual_toxic_lexicon.texttoken-classification100K<n<1M9 likes367 downloads2y agoHugging Face