datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FalaBracarense_splitsdataset website: projectofalabracarense
Licence
CC - BY - NC - ND
Restrictions: Academic - Non Commercial Use, Attribution, No Derivatives
barranquenho-ipa-dict-synthetic
Barranquenho IPA Pronunciation Dictionary
The first and only IPA pronunciation dictionary of Barranquenho — the
Ibero-Romance contact variety spoken in Barrancos (Baixo Alentejo, Portugal),
a mixed system born of centuries of Portuguese–Spanish (Extremaduran /
Andalusian) contact on the raia. Every headword is written in the
Convenção Ortográfica do Barranquenho (2025) orthography and paired with a
broad-phonemic IPA transcription plus Portuguese and Spanish glosses.
This… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/barranquenho-ipa-dict-synthetic.tigrinya-asr-merged
tigrinya-asr-merged
A merged Tigrinya speech-recognition dataset, combining and deduplicating:
badrex/tigrinya-speech (train pool)
google/WaxalNLP config tir_asr (train pool)
UBC-NLP/SimbaBench_dataset config asr_test_tir (held-out benchmark test set)
Processing
Standardized to audio (16kHz mono) and text columns, with a source column tracking origin
Unicode NFC-normalized transcripts, empty transcripts dropped
Exact-duplicate transcripts removed from the train… See the full description on the dataset page: https://huggingface.co/datasets/Harbidel/tigrinya-asr-merged.tigrinya-speech-data
Tigrinya Speech Data (Pooled)
A ~180.0-hour Tigrinya speech corpus, drawn from a single source (Afrivoice
Ethiopia) and filtered to only genuinely transcribed audio. Part of the
AfroNet multi-language TTS data
effort — sibling release to Yoruba/Hausa/Igbo/Kinyarwanda/Swahili, but Tigrinya (and
its four sibling Ethiopian-language releases, Amharic/Oromo/Sidama/Wolaytta) are
each published independently, not bundled into one combined "Ethiopia" dataset,
even though they share a… See the full description on the dataset page: https://huggingface.co/datasets/Professor/tigrinya-speech-data.SpokenPortugueseGeographicalSocialVarieties
Spoken Portuguese - Geographical and Social Varieties
dataset source: https://www.clul.ulisboa.pt
(1995-1997 - European Commission DGXXII, Programme LINGUA/SOCRATES)
The project is concluded and the materials are published in CD-ROM, with the exclusive publishing support of Instituto Camões, under the title Português Falado - Documentos Autênticos: Gravações áudio com transcrição alinhada. Its distribution outside of Portugal is ensured by Instituto Camões and in Portugal by CLUL.… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/SpokenPortugueseGeographicalSocialVarieties.ArquivoDialetalCLUPdataset info: https://cl.up.pt/arquivo/
license CC BY-NC-ND
VocativesEuropeanPortuguesedataset from https://www.clul.ulisboa.pt/en/recurso/vocatives-european-portuguese
This corpus was originally a corpus created for a study concerning with vocatives in European Portuguese. The main goal of this study was to analyze some prosodic features of the vocative in European Portuguese and their relation with the syntactic distribution (initial, medial, final) of these constituents.
The corpus has 432 audio files. This number results from the recording of 108 sentences (54 target… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/VocativesEuropeanPortuguese.leyu-tigrinya-speech-corpus-2026
Leyu Tigrinya Speech Corpus 2026
Official speech dataset submission for the Leyu Data Collection Competition 2026.
Organization & Team
Hugging Face Org: SoundWaveET
Dataset Repo: SoundWaveET/leyu-tigrinya-speech-corpus-2026
InstitutoCamoesdownloaded from https://www.instituto-camoes.pt
locallingua_ptRecordings from Portugal downloaded from https://localingual.com
compare-accents-ptsmall dataset of multiple portuguese speakers from various dialects speaking the same sentence
"Dom Sebastião I era o décimo-sexto Rei de Portugal, e sétimo da Dinastia de Avis. Era neto do rei João III, tornou-se herdeiro do trono depois da morte do seu pai, o príncipe João de Portugal duas semanas antes do seu nascimento, e rei com apenas três anos, em 1557. Em virtude de ser um herdeiro tão esperado para dar continuidade à Dinastia de Avis, ficou conhecido como O Desejado; alternativamente… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/compare-accents-pt.speech_MASSIVE_pt-PTpt-PT subset from FBK-MT/Speech-MASSIVE
pt_basicsphonetically diverse standalone words, letters, diphtongs and basic greetings
tigre-hubert-speech
Tigre HuBERT Speech Resources
Self-supervised speech resources for Tigre (ISO 639-3: tig), a Semitic
language spoken primarily in Eritrea and Sudan with very limited existing
speech-technology support. This repository bundles a Tigre-pretrained HuBERT
encoder, a discrete unit-discovery model, forced-aligned transcripts with
word-level unit sequences, and a word-to-unit pseudo-lexicon -- everything
needed to reproduce or extend this work.
Dataset Summary
6777… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-hubert-speech.
