CoolFace
Datasetpublic

mamed0v/TurkmenSpeech

Turkmen Speech Dataset (ASR) This dataset contains 251 hours of Turkmen speech audio with transcriptions, intended for training and evaluating Automatic Speech Recognition (ASR) models. It is one of the largest publicly available Turkmen speech datasets. Dataset Overview Property Value Total clips 119,847 Total duration 251.86 hours Sampling rate 16,000 Hz Language Turkmen (tk) Split train Each item includes: audio: waveform + sampling… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/TurkmenSpeech.

sourceHugging Facecc-by-nc-4.0updated 11mo agoView on Hugging Face
7likes1.4kdownloads
Dataset Card

Turkmen Speech Dataset (ASR)

This dataset contains 251 hours of Turkmen speech audio with transcriptions, intended for training and evaluating Automatic Speech Recognition (ASR) models. It is one of the largest publicly available Turkmen speech datasets.

Dataset Overview

PropertyValue
Total clips119,847
Total duration251.86 hours
Sampling rate16,000 Hz
LanguageTurkmen (tk)
Splittrain

Each item includes:

  • —audio: waveform + sampling rate
  • —text: Turkmen transcription
  • —duration: clip length in seconds
  • —start/end timestamps from the original source
  • —source: YouTube link
  • —speaker_id: currently unknown

Quality Notes

  • —95% of clips are under 20 seconds, suitable for ASR model training
  • —Contains real-world spoken Turkmen (news, interviews, dialogues, narration)
  • —Correct diacritic usage
  • —No empty or silent clips

License

This dataset is released under the Creative Commons Attribution–NonCommercial 4.0 International (CC BY-NC 4.0) license.

You are free to:

  • —Use, share, and modify the dataset for research and non-commercial purposes only

You must:

  • —Provide proper attribution
  • —Not use the dataset or derivative works for commercial purposes

Full license text: https://creativecommons.org/licenses/by-nc/4.0/

mamed0v/TurkmenSpeech · CoolFace