datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Tibetan-0310tibetan_voiceTibetanVoice: 6.5 hours of validated transcribed speech data from 9 audio book in lhasa dialect. The dataset is in tsv format with two columns, path and sentence. The path column contains the path to the audio file and the sentence column contains the corresponding sentence spoken in the audio file.tibetan-speech-english-text-dataset-new-updatedtibetan-english-8s_extend_speechtibetan-speech-english-text-dataset
Tibetan Speech Dataset with English Translations
Dataset Description
This dataset contains Tibetan speech recordings paired with transcriptions in Tibetan script and English translations. It is designed to support automatic speech recognition (ASR), machine translation, and text-to-speech (TTS) research for the Tibetan language, which is considered a low-resource language in NLP.
Supported Tasks
Automatic Speech Recognition (ASR): Train models to transcribe… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/tibetan-speech-english-text-dataset.tibetan_wz_ttsaugmented-merged-and-shuffel-tibetan-dataset
Augmented Audio Dataset
This dataset contains augmented audio samples with pitch shifting and white noise enhancement for improved model training and robustness.
Dataset Description
This is an augmented version of the original audio dataset, processed with audio augmentation techniques to increase dataset diversity and improve model generalization.
Augmentation Techniques Applied
Pitch Shifting: Random pitch shift between -2 to +2 semitones
White Noise… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/augmented-merged-and-shuffel-tibetan-dataset.tibetan-chinese-english-speech-augmented-3pipelinestibetan-to-english-audio-dataset
Tibetan to English Audio Dataset
Dataset Description
A Tibetan speech recognition dataset with transcriptions and English translations containing 1,642 audio samples.
Dataset Summary
This dataset contains Tibetan speech recordings with:
Tibetan transcriptions in native script
English translations
High-quality audio files in WAV format
Total Samples: 1,178Total Size: ~1.1 GBAudio Format: WAV
Languages
Source Language: Tibetan (བོད་སྐད་)
Target… See the full description on the dataset page: https://huggingface.co/datasets/Titung/tibetan-to-english-audio-dataset.tibetan-audio-to-english-fixed-filtered
Tibetan audio translation Dataset
Dataset Description
Tibetan audio translation Dataset
Dataset Summary
This dataset contains 6,366 audio samples with corresponding transcriptions, totaling approximately 15.8 hours of audio.
Languages
The dataset is in EN (Language code: en).
Dataset Structure
Data Fields
audio: An audio object containing:
path: Path to the audio file (if applicable)
array: Audio waveform as a numpy array… See the full description on the dataset page: https://huggingface.co/datasets/Titung/tibetan-audio-to-english-fixed-filtered.filtered-long-slowed-tibetan-datasetmerged-and-shuffel-tibetan-datasettibetan-chinese-english-speechmerged-tibetan-titung-goosetibetan-audio-english-6datasets-sample2
Tibetan Audio-English Sentence Dataset (Sample)
This is a sample dataset containing 5 rows from a merged collection of 6 Tibetan audio datasets with English translations.
📊 Dataset Details
Total Samples in Full Dataset: 17,278
Samples in This Preview: 5
Format: Audio + English sentence pairs
Audio Sampling Rate: 16,000 Hz
Languages: Tibetan (audio) → English (text)
🗂️ Source Datasets
This sample is merged from 6 datasets:… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/tibetan-audio-english-6datasets-sample2.tibetan_audio_data
Part 1
Dialect
Audio Count
Total Duration(hours)
Speaker Count
Kham(康巴)-Changdu(昌都)
2558
2.79
7
Kham(康巴)-Dege(德格)
1245
2.31
3
Kham(康巴)-Yushu(玉树)
631
0.77
3
Amdo(安多)-Farming-pastoral
3549
4.12
2
Amdo(安多)-Pastoral
19305
21.81
21
Tsang(藏区)-Shigatse(日喀则)10729
15.15
4
Tsang(藏区)-Lhasa(拉萨)
30349
37.38
48
tibetan_audio_english_textfiltered-long-tibetan-datasettibetan_plant_audio_castedtibetan-chinese-english-speech-augmentedtibetan-audio-english-6datasets-fullzangskari_Tibetan_test_predictionstibetan_plant_audio_casted_augmentedtibetan-audio-english-6datasets-sample
Tibetan Audio-English Sentence Dataset (Sample)
This is a sample dataset containing 5 rows from a merged collection of 6 Tibetan audio datasets with English translations.
📊 Dataset Details
Total Samples in Full Dataset: 17,278
Samples in This Preview: 5
Format: Audio + English sentence pairs
Audio Sampling Rate: 16,000 Hz
Languages: Tibetan (audio) → English (text)
🗂️ Source Datasets
This sample is merged from 6 datasets:… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/tibetan-audio-english-6datasets-sample.tibetan-audio-english-6datasets-sample1augmented_2_speed_pitch-merged-and-shuffel-tibetan-datasettibetan-english-8s_extend_audio_sentence_speech
