datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tibetan-speech-english-text-dataset
Tibetan Speech Dataset with English Translations
Dataset Description
This dataset contains Tibetan speech recordings paired with transcriptions in Tibetan script and English translations. It is designed to support automatic speech recognition (ASR), machine translation, and text-to-speech (TTS) research for the Tibetan language, which is considered a low-resource language in NLP.
Supported Tasks
Automatic Speech Recognition (ASR): Train models to transcribe… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/tibetan-speech-english-text-dataset.100_shruk-speech_to_Text__ASR_dataset
100 Speech-to-Text / ASR Shruk Dataset
Dataset Description
Traditional Kashmiri poetic verses (Shruks) paired with their Kashmiri script
transcriptions. Designed for training and evaluating Automatic Speech
Recognition (ASR) / Speech-to-Text (STT) models for the Kashmiri language.
Dataset Structure
Field
Type
Description
audio
Audio
WAV recording of the Shruk
transcription
string
Kashmiri script transcription
shruk_number
int
Original shruk… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/100_shruk-speech_to_Text__ASR_dataset.
