CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TigreGotico /arabic-stem-lexicon Arabic Diacritized-Stem Lexicon An undiacritized Arabic surface form → its most frequent diacritized stem. Standard Arabic writes no short vowels, so anything that has to pronounce Arabic must first put them back. A neural diacritizer does that well on rare words, where inference is the only thing there is. On common words it is the wrong tool: which vowels كتاب carries is not a thing to be inferred, it is a thing to be looked up — and models get exactly these wrong, reading… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-stem-lexicon.tabulartext-to-speech100K<n<1M0 likes3.1k downloads2mo agoHugging Face02beshiribrahim /tigre-hubert-dataaudio1K<n<10K0 likes816 downloads2mo agoHugging Face03TigreGotico /barranquenho-ipa-dict-synthetic Barranquenho IPA Pronunciation Dictionary The first and only IPA pronunciation dictionary of Barranquenho — the Ibero-Romance contact variety spoken in Barrancos (Baixo Alentejo, Portugal), a mixed system born of centuries of Portuguese–Spanish (Extremaduran / Andalusian) contact on the raia. Every headword is written in the Convenção Ortográfica do Barranquenho (2025) orthography and paired with a broad-phonemic IPA transcription plus Portuguese and Spanish glosses. This… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/barranquenho-ipa-dict-synthetic.texttext-to-speech1K<n<10K0 likes450 downloads2mo agoHugging Face04TigreGotico /mirandese_g2ptextn<1K1 likes424 downloads1y agoHugging Face05TigreGotico /ArquivoDialetalCLUP_ipadataset info: https://cl.up.pt/arquivo/ textn<1K0 likes416 downloads1y agoHugging Face06TigreGotico /portuguese-dialects-ipa-synthetic portuguese-dialects-ipa-synthetic 920 dialect-register sentences — 20 per lect across 46 lects: 41 Portuguese varieties (European regional, insular, Brazilian regional, African/Asian/border national norms, medieval stages) plus the other languages of Portugal and their kin: Mirandese (3 lects), Asturleonese of Portugal (Rionorese, Guadramilese) with Barranquenho, and Galician-Portuguese. Each row carries two IPA columns with distinct provenance. Schema sentence —… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-dialects-ipa-synthetic.texttext-to-speechn<1K0 likes122 downloads2mo agoHugging Face07TigreGotico /EAT EAT: Expected Answer Type Dataset A high-quality dataset for Question Classification based on the TREC Question Taxonomy, enhanced with modern categories and strict Expected Answer Type (EAT) validation. Dataset Summary The EAT (Expected Answer Type) dataset is designed to train and evaluate NLP models in the task of classifying questions not by their surface keywords, but by the semantic category of their expected answer. Unlike original TREC datasets, this version… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/EAT.texttext-classification10K<n<100K0 likes84 downloads5mo agoHugging Face08TigreGotico /dicionario_barranquenho Dicionário de Barranquenho Structured lexical dataset derived from the first published dictionary of Barranquenho, a Romance contact language spoken in Barrancos, Portugal. Contains 1,680 entries with Portuguese and Spanish glosses, grammatical categories, semantic fields, source attributions, and synonym cross-references. Language Barranquenho (glottocode: barr1245; no ISO 639-3 code assigned at time of publication) is a contact language spoken in the municipality of… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/dicionario_barranquenho.documenttranslation1K<n<10K0 likes52 downloads5mo agoHugging Face09TigreGotico /sentence-types-multilingual Little Questions: Multilingual Sentence Types Dataset A multilingual dataset of 69,300 labeled sentences (9,900 per language) across 6 sentence type categories and 7 languages. Designed for training and evaluating sentence-type classifiers in multilingual contexts. Dataset Details Total entries: 69,300 (9,900 × 7 languages) Languages: English (EN), Spanish (ES), French (FR), German (DE), Italian (IT), Portuguese (PT), Dutch (NL) Class distribution: 13,200 entries per… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/sentence-types-multilingual.texttext-classification10K<n<100K0 likes38 downloads6mo agoHugging Face10TigreGotico /yes-no-multilingual Yes/No Multilingual Answers Dataset A dataset of 8,600 conversational utterances for classifying yes/no/ambiguous responses across 43 languages. Dataset Description Each sample is a natural language utterance a person might say in response to a yes/no question. The dataset covers three classes: Label Description yes Affirmation, agreement, or confirmation no Negation, refusal, or disagreement None Genuinely ambiguous — cannot be resolved without context… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/yes-no-multilingual.texttext-classification1K<n<10K0 likes34 downloads5mo agoHugging Face11TigreGotico /portuguese_phonetic_lexicon 📚 Portuguese Phonetic Lexicon Dataset This dataset contains phonetic and morphological information for Portuguese words, collected from the Portal da Língua Portuguesa. It was generated by scraping the site across multiple Portuguese-speaking regions and dialects. 🌍 Regional Coverage The dataset includes words as spoken in ten regional variants: 🇵🇹 Lisbon (Standard and Non-Standard) 🇦🇴 Luanda 🇧🇷 Rio de Janeiro (Standard and Non-Standard) 🇧🇷 São Paulo (Standard… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese_phonetic_lexicon.texttext-classification100K<n<1M0 likes27 downloads7mo agoHugging Face12TigreGotico /arabic-mantoq-synthetic-g2ptexttext-generation1M<n<10M0 likes18 downloads9mo agoHugging Face13TigreGotico /chronologia-benchmark Chronologia temporal-extraction benchmark 1,049 natural-language temporal expressions with hand-derived gold dates, across 24 languages, scored against three parsers. Every gold value was derived by a human or by independent date arithmetic — never by any engine under test. Each row is one utterance a person might actually say ("3 semanas dimpués de o 15 de chinero de 2019", "the week of july 20 2026", "עד יום שישי") with: column meaning lang BCP-47 primary language… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/chronologia-benchmark.texttoken-classification1K<n<10K0 likes18 downloads2mo agoHugging Face14TigreGotico /sentence-types COMMAND: This label refers to statements that give instructions or orders. These can be further divided into sub-labels ACTION and DENIAL. ACTION: This sub-label refers to commands that instruct the listener to perform an action. DENIAL: This sub-label refers to commands that instruct the listener to refrain from performing an action. QUESTION: This label refers to statements that ask for information. These can be further divided into sub-labels QUERY, YESNO and REQUEST. QUERY: This… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/sentence-types.texttext-classification1K<n<10K0 likes14 downloads2y agoHugging Face15TigreGotico /atc-role-classification-synthetic-2procedurally generated synthetic data, script is included in files text100K<n<1M0 likes13 downloads7mo agoHugging Face16TigreGotico /atc-role-classification-synthetic-3Generated with ChatGPT text1K<n<10K0 likes9 downloads7mo agoHugging Face17TigreGotico /infopedia-pt-heterophones European Portuguese Heterophonic Homographs — Infopédia Full European-Portuguese dictionary data for a curated set of heterophonic homographs: words spelled identically but pronounced differently depending on the reading (e.g. sede ˈsɛdɨ "seat" vs ˈsedɨ "thirst"; jogo ˈʒoɡu "game" vs ˈʒɔɡu "I play"). The open/closed stressed-vowel contrast (ɔ/o, ɛ/e) is the core disambiguation axis for Portuguese grapheme-to-phoneme (G2P). One row per reading: each homograph contributes a row… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/infopedia-pt-heterophones.textn<1K0 likes9 downloads3mo agoHugging Face18TigreGotico /atc-role-classification-syntheticGenerated with Claude Opus 4.6 The data covers a wide range of realistic radio phraseology including: Pilot utterances: check-ins, readbacks, requests (altitude/heading/approach/deviation), position reports, PIREPs, emergencies (mayday/pan-pan), taxi/pushback, go-arounds, and general acknowledgments. ATC utterances: clearances (takeoff/landing/approach), altitude/heading/speed instructions, traffic advisories, hold instructions, weather broadcasts, vectors, STAR/SID assignments, taxi… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/atc-role-classification-synthetic.text10K<n<100K0 likes8 downloads7mo agoHugging Face19TigreGotico /portuguese-sentences-synthetic-g2p Dataset Card for 'TigreGotico/portuguese_g2p' Dataset Description Dataset Summary TigreGotico/portuguese_g2p is a Grapheme-to-Phoneme (G2P) dataset for Portuguese, offering phonetic transcriptions for sentences across ten different regional variants. It is derived from the portuguese_phonetic_lexicon and is designed to aid in the development of robust Speech Recognition (ASR) and Text-to-Speech (TTS) models that account for dialectal variation in Portuguese.… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-sentences-synthetic-g2p.texttext-classification10K<n<100K0 likes7 downloads1y agoHugging Face20TigreGotico /AO1990_pt-BRtext1K<n<10K0 likes7 downloads7mo agoHugging Face21TigreGotico /galician_g2ptexttext-generation1K<n<10K0 likes6 downloads1y agoHugging Face22TigreGotico /AO1990_pt-PTtext1K<n<10K0 likes6 downloads7mo agoHugging Face23TigreGotico /archaisms_ptAté ao início do século XX, tanto em Portugal como no Brasil, seguia-se uma ortografia que, por regra, baseava-se nos étimos latino ou grego para escrever cada palavra textn<1K0 likes2 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.