turkmen
Datasets
All datasets matching “turkmen”TurkmenSpeech
Turkmen Speech Dataset (ASR)
This dataset contains 251 hours of Turkmen speech audio with transcriptions, intended for training and evaluating Automatic Speech Recognition (ASR) models.
It is one of the largest publicly available Turkmen speech datasets.
Dataset Overview
Property
Value
Total clips
119,847
Total duration
251.86 hours
Sampling rate
16,000 Hz
Language
Turkmen (tk)
Split
train
Each item includes:
audio: waveform + sampling… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/TurkmenSpeech.TurkmenSpeech
Turkmen Speech Dataset (ASR)
This dataset contains 251 hours of Turkmen speech audio with transcriptions, intended for training and evaluating Automatic Speech Recognition (ASR) models.
It is one of the largest publicly available Turkmen speech datasets.
Dataset Overview
Property
Value
Total clips
119,847
Total duration
251.86 hours
Sampling rate
16,000 Hz
Language
Turkmen (tk)
Split
train
Each item includes:
audio: waveform + sampling… See the full description on the dataset page: https://huggingface.co/datasets/rozumov/TurkmenSpeech.ipfs_turkmenistan_laws_ir
Turkmenistan legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_turkmenistan_laws (revision ce38a3085e088c63199c8df68f15274c86755d74) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Turkmenistan prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_turkmenistan_laws_ir.ipfs_turkmenistan_laws
Turkmenistan Free Mejlis Codes and Laws (mejlis.gov.tm)
Research snapshot of official national legislation from Mejlis of Turkmenistan free HTML (mejlis.gov.tm); Adalat paid DB skipped.
Not legal advice. The official gazette / authentic source prevails over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-09
Coverage
free-mejlis-codes-laws-batch
Source
Mejlis of Turkmenistan free HTML (mejlis.gov.tm); Adalat paid DB skipped
Collector… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_turkmenistan_laws.alpaca-turkmen
Turkmen Alpaca Dataset
Overview
This dataset is a Turkmen translation of the original Alpaca dataset. The Alpaca dataset is a publicly available instruction-following dataset containing approximately 52,000 instruction-following samples. This Turkmen version aims to extend the accessibility of instruction-following datasets to the Turkmen language community.
Dataset Details
Original Dataset: Alpaca
Languages: English and Turkmen
Number of Samples:… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/alpaca-turkmen.orca-math-word-problems-200k-turkmen
Turkmen Orca Math Word Problems 200k Dataset
Overview
This dataset is a Turkmen translation of the original microsoft/orca-math-word-problems-200k dataset. The Orca Math Word Problems dataset contains 200,000 high-quality math word problems and their solutions. This Turkmen version aims to extend the accessibility of math problem-solving datasets to the Turkmen language community.
Dataset Details
Original Dataset: microsoft/orca-math-word-problems-200k… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/orca-math-word-problems-200k-turkmen.
