datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
uyghur-common-voice-tts
Uyghur Common Voice TTS Dataset
A cleaned and processed Text-to-Speech (TTS) dataset for the Uyghur language, derived from Mozilla Common Voice.
Dataset Summary
Property
Value
Language
Uyghur (ug)
Total Samples
43,054
Train Samples
40,901
Validation Samples
2,153
Audio Format
WAV
Source
Mozilla Common Voice
License
CC0-1.0
Dataset Structure
/
├── train.jsonl # Training data (40,901 samples)
├── val.jsonl #… See the full description on the dataset page: https://huggingface.co/datasets/anke01/uyghur-common-voice-tts.common_voice_25_0-word_alignmentCommonVoice-SpeechRE-textThis repository provides the text annotation part of the CommonVoice-SpeechRE dataset, a benchmark for Speech Relation Extraction (SpeechRE).It includes transcripts, entity annotations, and relation labels aligned with speech IDs.
👉 The corresponding audio subset (19,583 samples, downsampled to 16kHz) is available at:CommonVoice-SpeechRE-audio
Dataset Details
Source transcripts: From Common Voice 17.0
Annotations: Entities and relations manually labeled by our team… See the full description on the dataset page: https://huggingface.co/datasets/DMU-ITREC/CommonVoice-SpeechRE-text.quantized-common-voice-encommon-voice-en-revoicetelugu-commonvoice-tts-wavcommon-voice-word-alignments-testcommon_voice_25_0-word_alignment-testMozila-Common-Voice-page-Emakhuwa-localizationtelugu-commonvoice-tts
