datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TibetanSft_corpusTibetanGeneral_corpus
此数据为网上收集的藏语单语数据集,规模为258661条,经过预处理以及清洗,可用于预训练。
数据格式如下所示:
{
"taskname": "用于预训练的单语数据集",
"url": "",
"instruction": "公开数据集",
"input": "ཚན་རིག་ནི་དང་ཐོག་རང་བྱུང་ཁྱབ་ཁོངས་ཀྱི་ཤེས་བྱ་ཡིན་ཞིང་འདི་ནས་སྤྱི་ཚོགས་དང་བསམ་བློ་ལ་སོགས་སུ་ཁྱབ་ཆེ་རུ་ཕྱིན།དཔེར་ནི་སྤྱི་ཚོགས་ཚན་རིག་ལྟ་བུ།",
"output":""
}
tibetan-pretraining-corpus
Tibetan Pre-training Corpus
This dataset provides a comprehensive Tibetan language corpus developed for pre-training large language models. It contains carefully curated text from diverse sources to ensure broad coverage of the Tibetan language.
Dataset Description
The corpus consists of approximately 500,000 words of Tibetan text collected and processed specifically for large language model adaptation. The data has been sourced from:
Tibetan Wikipedia content (70% of… See the full description on the dataset page: https://huggingface.co/datasets/lightman7/tibetan-pretraining-corpus.tibetan-mix-instruction-tuning-60K
Tibetan Instruction Tuning Dataset
This dataset provides instruction-response pairs in the Tibetan language specifically designed for instruction fine-tuning of large language models.
Dataset Description
The dataset contains 60,000 instruction-response pairs mix with Chinese and Tibetan, derived from:
Translated alpaca-gpt4 dataset (English instruction-following responses generated by GPT-4)
Chinese-Tibetan translation data
All content has undergone quality filtering to… See the full description on the dataset page: https://huggingface.co/datasets/lightman7/tibetan-mix-instruction-tuning-60K.
