CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TigreGotico /FalaBracarense_splitsdataset website: projectofalabracarense Licence CC - BY - NC - ND Restrictions: Academic - Non Commercial Use, Attribution, No Derivatives audioautomatic-speech-recognition100K<n<1M0 likes1.6k downloads1y agoHugging Face02TigreGotico /barranquenho-ipa-dict-synthetic Barranquenho IPA Pronunciation Dictionary The first and only IPA pronunciation dictionary of Barranquenho — the Ibero-Romance contact variety spoken in Barrancos (Baixo Alentejo, Portugal), a mixed system born of centuries of Portuguese–Spanish (Extremaduran / Andalusian) contact on the raia. Every headword is written in the Convenção Ortográfica do Barranquenho (2025) orthography and paired with a broad-phonemic IPA transcription plus Portuguese and Spanish glosses. This… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/barranquenho-ipa-dict-synthetic.texttext-to-speech1K<n<10K0 likes592 downloads2mo agoHugging Face03Harbidel /tigrinya-asr-mergedgated tigrinya-asr-merged A merged Tigrinya speech-recognition dataset, combining and deduplicating: badrex/tigrinya-speech (train pool) google/WaxalNLP config tir_asr (train pool) UBC-NLP/SimbaBench_dataset config asr_test_tir (held-out benchmark test set) Processing Standardized to audio (16kHz mono) and text columns, with a source column tracking origin Unicode NFC-normalized transcripts, empty transcripts dropped Exact-duplicate transcripts removed from the train… See the full description on the dataset page: https://huggingface.co/datasets/Harbidel/tigrinya-asr-merged.audioautomatic-speech-recognition10K<n<100K0 likes73 downloads22d agoHugging Face04Professor /tigrinya-speech-data Tigrinya Speech Data (Pooled) A ~180.0-hour Tigrinya speech corpus, drawn from a single source (Afrivoice Ethiopia) and filtered to only genuinely transcribed audio. Part of the AfroNet multi-language TTS data effort — sibling release to Yoruba/Hausa/Igbo/Kinyarwanda/Swahili, but Tigrinya (and its four sibling Ethiopian-language releases, Amharic/Oromo/Sidama/Wolaytta) are each published independently, not bundled into one combined "Ethiopia" dataset, even though they share a… See the full description on the dataset page: https://huggingface.co/datasets/Professor/tigrinya-speech-data.text-to-speech10K<n<100K0 likes70 downloads1mo agoHugging Face05TigreGotico /SpokenPortugueseGeographicalSocialVarieties Spoken Portuguese - Geographical and Social Varieties dataset source: https://www.clul.ulisboa.pt (1995-1997 - European Commission DGXXII, Programme LINGUA/SOCRATES) The project is concluded and the materials are published in CD-ROM, with the exclusive publishing support of Instituto Camões, under the title Português Falado - Documentos Autênticos: Gravações áudio com transcrição alinhada. Its distribution outside of Portugal is ensured by Instituto Camões and in Portugal by CLUL.… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/SpokenPortugueseGeographicalSocialVarieties.audioautomatic-speech-recognitionn<1K0 likes59 downloads1y agoHugging Face06TigreGotico /ArquivoDialetalCLUPdataset info: https://cl.up.pt/arquivo/ license CC BY-NC-ND audioautomatic-speech-recognitionn<1K0 likes56 downloads1y agoHugging Face07TigreGotico /VocativesEuropeanPortuguesedataset from https://www.clul.ulisboa.pt/en/recurso/vocatives-european-portuguese This corpus was originally a corpus created for a study concerning with vocatives in European Portuguese. The main goal of this study was to analyze some prosodic features of the vocative in European Portuguese and their relation with the syntactic distribution (initial, medial, final) of these constituents. The corpus has 432 audio files. This number results from the recording of 108 sentences (54 target… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/VocativesEuropeanPortuguese.audioautomatic-speech-recognitionn<1K1 likes32 downloads1y agoHugging Face08SoundWaveET /leyu-tigrinya-speech-corpus-2026 Leyu Tigrinya Speech Corpus 2026 Official speech dataset submission for the Leyu Data Collection Competition 2026. Organization & Team Hugging Face Org: SoundWaveET Dataset Repo: SoundWaveET/leyu-tigrinya-speech-corpus-2026 audioautomatic-speech-recognitionn<1K0 likes31 downloads1mo agoHugging Face09TigreGotico /InstitutoCamoesdownloaded from https://www.instituto-camoes.pt audioautomatic-speech-recognitionn<1K1 likes26 downloads1y agoHugging Face10TigreGotico /locallingua_ptRecordings from Portugal downloaded from https://localingual.com audioautomatic-speech-recognitionn<1K0 likes25 downloads2y agoHugging Face11TigreGotico /compare-accents-ptsmall dataset of multiple portuguese speakers from various dialects speaking the same sentence "Dom Sebastião I era o décimo-sexto Rei de Portugal, e sétimo da Dinastia de Avis. Era neto do rei João III, tornou-se herdeiro do trono depois da morte do seu pai, o príncipe João de Portugal duas semanas antes do seu nascimento, e rei com apenas três anos, em 1557. Em virtude de ser um herdeiro tão esperado para dar continuidade à Dinastia de Avis, ficou conhecido como O Desejado; alternativamente… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/compare-accents-pt.audioautomatic-speech-recognitionn<1K1 likes20 downloads2y agoHugging Face12TigreGotico /speech_MASSIVE_pt-PTpt-PT subset from FBK-MT/Speech-MASSIVE audioautomatic-speech-recognition1K<n<10K1 likes18 downloads1y agoHugging Face13TigreGotico /pt_basicsphonetically diverse standalone words, letters, diphtongs and basic greetings audioautomatic-speech-recognitionn<1K0 likes15 downloads1y agoHugging Face14BeitTigreAI /tigre-hubert-speechgated Tigre HuBERT Speech Resources Self-supervised speech resources for Tigre (ISO 639-3: tig), a Semitic language spoken primarily in Eritrea and Sudan with very limited existing speech-technology support. This repository bundles a Tigre-pretrained HuBERT encoder, a discrete unit-discovery model, forced-aligned transcripts with word-level unit sequences, and a word-to-unit pseudo-lexicon -- everything needed to reproduce or extend this work. Dataset Summary 6777… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-hubert-speech.audioautomatic-speech-recognition1K<n<10K0 likes12 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.