datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
uyghur-dictionary-dataset
维吾尔语多语言词典数据集
维吾尔语-汉语-英语多语言词典数据集,适用于大型语言模型(LLM)微调训练。
数据集统计
数据集
条目数
大小
语言方向
ug-cn.jsonl
3,079,016
456 MB
维吾尔语 ⟷ 汉语
ug-en.jsonl
685,836
108 MB
维吾尔语 ⟷ 英语
en-ug.jsonl
705,534
112 MB
英语 ⟷ 维吾尔语
ug-ug.jsonl
139,920
46 MB
维吾尔语释义
cn-cn.jsonl
127,646
39 MB
汉语释义
总计: 4,737,952 条(双向)/ 约 236 万唯一对
数据格式
{
"instruction": "请翻译以下维吾尔语词汇",
"input": "مەركىزى",
"output": "中心的,中央的"
}
字段说明:
instruction: 任务指令
input: 输入文本
output: 输出文本
使用示例… See the full description on the dataset page: https://huggingface.co/datasets/anke01/uyghur-dictionary-dataset.uyghur-ASR-dataset
Uyghur ASR Corpus (Latin Transliteration)
A speech corpus for Uyghur automatic speech recognition, with transcriptions in a
case-sensitive Latin transliteration scheme. Approximately 23 hours of audio across
9,468 clips.
Uyghur is a Turkic language spoken by roughly 10–12 million people. It is severely
under-represented in open speech datasets, and this corpus is intended to support ASR research
for the language.
Dataset summary
Language
Uyghur (ug)… See the full description on the dataset page: https://huggingface.co/datasets/Shramadeepd/uyghur-ASR-dataset.uyghur_ner_dataset
Uyghur NER dataset
Description
This dataset is in WikiAnn format. The dataset is assembled from named entities parsed from Wikipedia, Wiktionary and Dbpedia. For some words, new case forms have been created using Apertium-uig. Some locations have been translated using the Google Translate API.
The dataset is divided into two parts: train and extra. Train has full sentences, extra has only named entities.
Tags: O (0), B-PER (1), I-PER (2), B-ORG (3), I-ORG (4), B-LOC (5)… See the full description on the dataset page: https://huggingface.co/datasets/codemurt/uyghur_ner_dataset.uyghur-asr-ranking-dataset
