CoolFace
20 results

tigre

TigreGotico /arabic-stem-lexicon Arabic Diacritized-Stem Lexicon An undiacritized Arabic surface form → its most frequent diacritized stem. Standard Arabic writes no short vowels, so anything that has to pronounce Arabic must first put them back. A neural diacritizer does that well on rare words, where inference is the only thing there is. On common words it is the wrong tool: which vowels كتاب carries is not a thing to be inferred, it is a thing to be looked up — and models get exactly these wrong, reading… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-stem-lexicon.tabulartext-to-speech100K<n<1M0 likes3.1k downloads2mo agoHugging FaceTigreGotico /FalaBracarense_splitsdataset website: projectofalabracarense Licence CC - BY - NC - ND Restrictions: Academic - Non Commercial Use, Attribution, No Derivatives audioautomatic-speech-recognition100K<n<1M0 likes2.2k downloads1y agoHugging Facebeshiribrahim /tigre-hubert-dataaudio1K<n<10K0 likes816 downloads2mo agoHugging FaceTigreGotico /arabic-dialects-gold20 arabic-dialects-gold20 660 sentences: 33 Arabic lects × 20 sentences, each with fully diacritized dialectal orthography, the undiacritized surface form, gold IPA, an engine draft, an English gloss, machine-verified phonetic feature tags, per-row verification metadata, and notes citing the dialectological literature that grounds the row. Columns (TSV, UTF-8, one file per lect): id, sentence, raw, ipa, ipa_o2i, gloss_en, features, notes, fable_corrections, verification… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-dialects-gold20.texttext-to-speechn<1K0 likes758 downloads2mo agoHugging FaceTigreGotico /not-wake-words-speech-en not-wake-words-speech-en Negative (non-wake-word) speech clips, used to measure false accepts for OVOS wake-word plugins. Derived from the Multilingual Spoken Words Corpus (MLCommons), which is built from Mozilla Common Voice and licensed CC-BY-4.0. This derivative keeps the same licence and attribution requirement. Produced with support from the NGI0 Commons Fund. audioaudio-classification10K<n<100K0 likes584 downloads21d agoHugging FaceTigreGotico /portuguese-unified-pronunciation-lexicon Portuguese Unified Pronunciation Lexicon A flat, single-row-per-pronunciation dataset merging Portuguese IPA transcriptions from three authoritative sources. Each row is a word × region × POS tuple with both broad phonemic (ipa_broad) and narrow phonetic (ipa_narrow) transcriptions normalized across sources. Source Words Convention Description Infopédia (Porto Editora) 102,685 Broad phonemic European Portuguese dictionary IPA Wiktionary (pt.wiktionary.org) 15,720… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-unified-pronunciation-lexicon.texttext-generation100K<n<1M1 likes480 downloads2mo agoHugging Face