datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tibetan-voice-benchmarkBenchmark of Tibetan Speech-To-Text dataset created by Monlam AI x Openpecha in 2024
All the transcripts have been reviewed by at least one person in addition to the original transcriber.
Data was taken on 15 July 2024 02∶47∶06 PM.
dept
desc
Count
STT_AB
Audio book
1000
STT_CS
Children Speech
1367
STT_HS
History
1000
STT_MV
Tibetan Movies
1000
STT_NS
Natural Speech
1000
STT_NW
News
1000
STT_PC
Podcast
1000
STT_TT
Tibetan Teachings
1000
grade column is used to… See the full description on the dataset page: https://huggingface.co/datasets/openpecha/tibetan-voice-benchmark.tibetan-to-spanish-translation-datasetThis dataset consists of three columns, the first of which is a sentence or phrase in Tibetan, the second is the phonetic transliteration of the Tibetan, and the third is the Spanish translation of the Tibetan.
The dataset was scraped from Lotsawa House and is released under the same license as the texts from which it is sourced.
The dataset is part of the larger MLotsawa project, the code repo for which can be found here.
tibetan_sa
Sentiment Analysis Data for the Tibetan Language
Dataset Description:
This dataset contains a sentiment analysis data from Zhu et al. (2023).
Data Structure:
The data was used for the project on injecting external commonsense knowledge into multilingual Large Language Models.
Citation:
@INPROCEEDINGS{10348366,
author={Zhu, Yulei and Luosai, Baima and Zhou, Liyuan and Qun, Nuo and Nyima, Tashi},
booktitle={2023 IEEE 4th International Conference on Pattern Recognition and Machine… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/tibetan_sa.tibetan-to-english-translation-datasetThis dataset consists of three columns, the first of which is a sentence or phrase in Tibetan, the second is the phonetic transliteration of the Tibetan, and the third is the English translation of the Tibetan.
This dataset is identical to 'billingsmoore/tibetan-phonetic-transliteration' with the addition of a third column of translations.
The dataset was scraped from Lotsawa House and is released under the same license as the texts from which it is sourced.
The dataset was used to train the… See the full description on the dataset page: https://huggingface.co/datasets/billingsmoore/tibetan-to-english-translation-dataset.tibetan-to-german-translation-datasetThis dataset consists of three columns, the first of which is a sentence or phrase in Tibetan, the second is the phonetic transliteration of the Tibetan, and the third is the German translation of the Tibetan.
The dataset was scraped from Lotsawa House and is released under the same license as the texts from which it is sourced.
The dataset is part of the larger MLotsawa project, the code repo for which can be found here.
tibetan-to-french-translation-datasetThis dataset consists of three columns, the first of which is a sentence or phrase in Tibetan, the second is the phonetic transliteration of the Tibetan, and the third is the French translation of the Tibetan.
The dataset was scraped from Lotsawa House and is released under the same license as the texts from which it is sourced.
The dataset is part of the larger MLotsawa project, the code repo for which can be found here.
tibetan_ranking_training_data
Tibetan Ranking Training Data — Tier 1
Training data for a cross-encoder that re-ranks Tibetan→English translation candidates by contextual relevance. Each row pairs a Tibetan term and sentence context with one English gloss candidate and a binary label: 1 (this gloss is defensible in this context) or 0 (it is not).
Used to train trabten/tibetan_ranking_BUDA.
Files
File
Rows
Description
training_v2_merged.csv
71,050
Tier 1 — surgically annotated… See the full description on the dataset page: https://huggingface.co/datasets/trabten/tibetan_ranking_training_data.phonetic-tibetan-to-english-translation-pairsThis dataset consists of a single csv with 10,000,000 rows which are pairs of sentences or phrases. The first member of each pair is a sentence or phrase in Classical Tibetan. The second member is the English translation of the first.
The pairs are pulled from texts sourced from Lotsawa House and are offered under the same license as the original texts from which they are sourced.
This data was scraped, cleaned, and formatted programmatically. Because of the difficulty in assembling data in… See the full description on the dataset page: https://huggingface.co/datasets/billingsmoore/phonetic-tibetan-to-english-translation-pairs.Tibetan_transliteration_wylie_datasetstibetan_conceptnet
ConceptNet Data for the Tibetan Language
Dataset Description:
This dataset contains data extracted from ConceptNet using the dedicated module for fetching knowledge from the graph, available on GitHub.
Data Structure:
The data is converted from triplets into natural text using a pre-defined relationship mapping and split into training and validation sets. It was used for training language adapters for the project aimed at injecting external commonsense knowledge into multilingual… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/tibetan_conceptnet.tibetan-phonetic-transliteration-datasetThis dataset has been superseded by 'billingsmoore/tibetan-to-english-translation-dataset'
This dataset consists of a single csv file with 98,597 rows. There are no column headers. The first column is a line of unicode Tibetan script. The second column is the phonetic transliteration of that Tibetan script.
The dataset was scraped from Lotsawa House and is released under the same license as the texts from which it is sourced.
The dataset was used to train the model… See the full description on the dataset page: https://huggingface.co/datasets/billingsmoore/tibetan-phonetic-transliteration-dataset.
