datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tigre-hubert-speech
Tigre HuBERT Speech Resources
Self-supervised speech resources for Tigre (ISO 639-3: tig), a Semitic
language spoken primarily in Eritrea and Sudan with very limited existing
speech-technology support. This repository bundles a Tigre-pretrained HuBERT
encoder, a discrete unit-discovery model, forced-aligned transcripts with
word-level unit sequences, and a word-to-unit pseudo-lexicon -- everything
needed to reproduce or extend this work.
Dataset Summary
6777… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-hubert-speech.hub5_english_eval_2000_swb1
