datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TurkmenSpeech
Turkmen Speech Dataset (ASR)
This dataset contains 251 hours of Turkmen speech audio with transcriptions, intended for training and evaluating Automatic Speech Recognition (ASR) models.
It is one of the largest publicly available Turkmen speech datasets.
Dataset Overview
Property
Value
Total clips
119,847
Total duration
251.86 hours
Sampling rate
16,000 Hz
Language
Turkmen (tk)
Split
train
Each item includes:
audio: waveform + sampling… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/TurkmenSpeech.TurkmenSpeech
Turkmen Speech Dataset (ASR)
This dataset contains 251 hours of Turkmen speech audio with transcriptions, intended for training and evaluating Automatic Speech Recognition (ASR) models.
It is one of the largest publicly available Turkmen speech datasets.
Dataset Overview
Property
Value
Total clips
119,847
Total duration
251.86 hours
Sampling rate
16,000 Hz
Language
Turkmen (tk)
Split
train
Each item includes:
audio: waveform + sampling… See the full description on the dataset page: https://huggingface.co/datasets/rozumov/TurkmenSpeech.ipfs_turkmenistan_laws_ir
Turkmenistan legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_turkmenistan_laws (revision ce38a3085e088c63199c8df68f15274c86755d74) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Turkmenistan prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_turkmenistan_laws_ir.ipfs_turkmenistan_laws
Turkmenistan Free Mejlis Codes and Laws (mejlis.gov.tm)
Research snapshot of official national legislation from Mejlis of Turkmenistan free HTML (mejlis.gov.tm); Adalat paid DB skipped.
Not legal advice. The official gazette / authentic source prevails over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-09
Coverage
free-mejlis-codes-laws-batch
Source
Mejlis of Turkmenistan free HTML (mejlis.gov.tm); Adalat paid DB skipped
Collector… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_turkmenistan_laws.alpaca-turkmen
Turkmen Alpaca Dataset
Overview
This dataset is a Turkmen translation of the original Alpaca dataset. The Alpaca dataset is a publicly available instruction-following dataset containing approximately 52,000 instruction-following samples. This Turkmen version aims to extend the accessibility of instruction-following datasets to the Turkmen language community.
Dataset Details
Original Dataset: Alpaca
Languages: English and Turkmen
Number of Samples:… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/alpaca-turkmen.orca-math-word-problems-200k-turkmen
Turkmen Orca Math Word Problems 200k Dataset
Overview
This dataset is a Turkmen translation of the original microsoft/orca-math-word-problems-200k dataset. The Orca Math Word Problems dataset contains 200,000 high-quality math word problems and their solutions. This Turkmen version aims to extend the accessibility of math problem-solving datasets to the Turkmen language community.
Dataset Details
Original Dataset: microsoft/orca-math-word-problems-200k… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/orca-math-word-problems-200k-turkmen.turkmen_english_s500
TL;DR & Quick results
Try it on Space demo Article with full technical journey is available Medium.
Dataset Description
(This dataset was created particularly to experiment with fine-tuning NLLB-200 model. You can try the outcome model on this Space or observe the model here)
Dataset Summary
This dataset provides a parallel corpus of small sentences in Turkmen (tk) and English (en). It consists of approximately 500-700 sentence pairs, designed primarily for machine… See the full description on the dataset page: https://huggingface.co/datasets/XSkills/turkmen_english_s500.alpaca-turkmen-cleanedThis repository contains the dataset used for the TaCo paper.
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation
@inproceedings{upadhayay2024taco,
title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes},
author={Bibek Upadhayay and Vahid Behzadan},
booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-turkmen-cleaned.turkmen-tts-cleanTurkmenTrilingualSemi-SyntheticDictionaryDF
🏜️ Turkmen Trilingual Semi-Synthetic — Dialogue Format
Language: Turkmen 🇹🇲 | English 🇬🇧 | Russian 🇷🇺Type: Instruction-style / Dialogue datasetRecords: 61 970 base recordsDialog turns (flattened): 378 941Splits: train=363 783, val=7 578, test=7 580
📘 Overview
This dataset is a dialogue-style extension of the original mamed0v/TurkmenTrilingualSemi-SyntheticDictionary.It was reformatted into conversational pairs to better suit instruction-tuning, chatbot… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/TurkmenTrilingualSemi-SyntheticDictionaryDF.generated-ud-turkmen
Data Splits
Split
Description
train
Automatically translated and silver-annotated sentences derived from UD training sources
dev
Silver-annotated evaluation data derived from UD test resources
test
Silver-annotated evaluation data derived from UD test resources
Dataset Creation
Source Data
This dataset is a derived work based on resources from:
Universal Dependencies treebanks - https://github.com/UniversalDependencies
Turkic UD… See the full description on the dataset page: https://huggingface.co/datasets/turkicnlp/generated-ud-turkmen.ahmetkayavocals
Ahmet Kaya Vokal
Dataset Açıklaması
Ahmet Kaya'nın şarkılarındaki vokaller UVR ile çıkarılmıştır. Toplam 103 adet veri vardır.
ahmetv3.zip içinde ise Ahmet Kayanın en popüler 3-4 röportajının kesilerek sadece onun konuştuğu kısımlar ayrıca UVR ile tekrar gözden geçirilerek oluşturulmuş hali vardır.
ak320kbps.zip'in içinde ise Ahmet Kaya'nın 320kbps olarak indirilmiş yüksek kalite olduğu iddia edilen 4 parçasının UVR windows size 1024 agression 15 1_HP_UVR modeli ile… See the full description on the dataset page: https://huggingface.co/datasets/turkmen/ahmetkayavocals.dolly-15k-turkmen
Turkmen Dolly 15k Dataset
Overview
This dataset is a Turkmen translation of the original Dolly 15k dataset. The Dolly dataset is a publicly available instruction-following dataset created by Databricks, containing 15,000 high-quality human-generated prompt-response pairs. This Turkmen version aims to extend the accessibility of instruction-following datasets to the Turkmen language community.
Dataset Details
Original Dataset: Dolly 15k
Language: Turkmen
Number… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/dolly-15k-turkmen.TurkmenTrilingualSemi-SyntheticDictionary
📚 Turkmen Trilingual Semi-Synthetic Dictionary
🌍 Обзор
Этот датасет содержит 61 970 триязычных словарных записей (туркменский–английский–русский), дополненных синтетически сгенерированными примерами использования. Заголовочные слова и их первоначальные переводы были извлечены из различных туркменских PDF-словарей, что делает датасет «полусинтетическим».
Языки: туркменский (tk), английский (en), русский (ru)
Формат: JSONL
Размер: 61 970 записей
Источник: 18… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/TurkmenTrilingualSemi-SyntheticDictionary.asia-aid-flows-iati-turkmenistan
Turkmenistan - Current IATI Aid Activities
Publisher: International Aid Transparency Initiative · Source: HDX · License: hdx-other · Updated: 2026-05-05
Abstract
List of active aid activities shared via the International Aid Transparency Initiative (IATI). Includes both humanitarian and development activities. More information on each activity (including financial data) is available from http://www.d-portal.org
Each row in this dataset represents country-level… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-aid-flows-iati-turkmenistan.asia-education-hdx-hapi-turkmenistan
HDX HAPI Data for Turkmenistan
Publisher: HDX Humanitarian API Data · Source: HDX · License: hdx-other · Updated: 2026-02-18
Abstract
This dataset contains data obtained from the
HDX Humanitarian API (HDX HAPI),
which provides standardized humanitarian indicators designed
for seamless interoperability from multiple sources.
The data facilitates automated workflows and visualizations
to support humanitarian decision making.
For more information, please see the HDX HAPI… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-education-hdx-hapi-turkmenistan.turkmen_hukukalpaca_turkmen_tacoThis repository contains the dataset used for the TaCo paper.
The dataset follows the style outlined in the TaCo paper, as follows:
{
"instruction": "instruction in xx",
"input": "input in xx",
"output": "Instruction in English: instruction in en ,
Response in English: response in en ,
Response in xx: response in xx "
}
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_turkmen_taco.asia-disability-who-data-for-turkmenistan
Turkmenistan - Health Indicators
This dataset contains data from WHO's data portal covering the following categories: Air pollution, Antimicrobial resistance (AMR), Assistive technology, Child mortality, Dementia diagnosis, treatment and care, Dementia policy and legislation, Environment and heal
Resources
Resource Count: 34
Formats: csv
Last Updated: 2025-02-07
Coverage
tkm
Tags
disability, disease, environment, health, hxl, indicators… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-disability-who-data-for-turkmenistan.asia-aviation-ourairports-turkmenistan
Airports in Turkmenistan
Publisher: OurAirports · Source: HDX · License: Public Domain · Updated: 2026-04-09
Abstract
List of airports in Turkmenistan, with latitude and longitude. Unverified community data from http://ourairports.com/countries/TM/
Each row in this dataset represents first-level administrative unit observations. Data was last updated on HDX on 2026-04-09. Geographic scope: TKM.
Curated into ML-ready Parquet format by Electric Sheep Africa.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-aviation-ourairports-turkmenistan.asia-ports-turkmenistan-daily-port-activity-data-an
Turkmenistan: Daily Port Activity Data and Shipment Estimates
Publisher: PortWatch · Source: HDX · License: hdx-other · Updated: 2026-05-06
Abstract
Daily count of port calls, estimates of incoming shipment volumes and outgoing shipment volumes (in metric tons) for ports in Turkmenistan.
Each row in this dataset represents country-level aggregates. Data was last updated on HDX on 2026-05-06. Geographic scope: TKM.
Curated into ML-ready Parquet format by Electric Sheep… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-ports-turkmenistan-daily-port-activity-data-an.turkmen-martyrs-dataset
🌍 Turkmen Martyrs Dataset / Türkmen Şehitleri Veri Seti / مجموعة بيانات شهداء التركمان
الرحمة والخلود لجميع شهدائنا الأبرار. / Dedicated to the eternal memory of our martyrs. / Şehitlerimizin aziz hatırasına ithafen.
يهدف هذا المستودع المفتوح المصدر إلى توثيق وحفظ السجلات التاريخية لشهداء التركمان ليكون مرجعاً موثوقاً للباحثين والمؤرخين، ولضمان عدم نسيان تضحياتهم.
📊 Dataset Structure (هيكل البيانات)
تتوفر البيانات بصيغة CSV وتحتوي على الأعمدة التالية:
id… See the full description on the dataset page: https://huggingface.co/datasets/aab20abdullah/turkmen-martyrs-dataset.asia-who-historical-data-for-turkmenistan
Turkmenistan - Historical Health Indicators
Publisher: World Health Organization · Source: HDX · License: hdx-other · Updated: 2025-02-07
Abstract
This dataset contains historical data from WHO's data portal.
Each row in this dataset represents first-level administrative unit observations. Data was last updated on HDX on 2025-02-07. Geographic scope: TKM.
Curated into ML-ready Parquet format by Electric Sheep Africa.
Dataset Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-who-historical-data-for-turkmenistan.turkmence_all_dataasia-demographics-dhs-data-for-turkmenistan
Turkmenistan - National Demographic and Health Data
Contains data from the DHS data portal. There is also a dataset containing Turkmenistan - Subnational Demographic and Health Data on HDX. The DHS Program Application Programming Interface (API)
Resources
Resource Count: 32
Formats: csv
Last Updated: 2026-04-20
Coverage
tkm
Tags
demographics, health
License
License ID: hdx-other
Citation
@dataset{{{slug}},
title =… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-demographics-dhs-data-for-turkmenistan.turkmen1xbet-turkmenistan-zerkalo-tmcell
1xBet Türkmenistan — Işleýän Aýna | Рабочее Зеркало 1xBet TMCell 2026
🇹🇲 1xBet Türkmenistanda nädip açmaly? (VPN-siz işleýän täze salgy)
1xBet saýty petiklenen bolsa, täze işjeň aýna (zerkalo) we TMCell / AGTS arkaly göni girmek üçin aşakdaky resmi düwmeleri ulanyň.
🚀 Göni saýta girmek: 1xBet Resmi Aýna (onex-tm.cc)
🤖 Telegram Aýna Bot: @onex_tm_bot
📱 Android APK: TMCell ulgamynda VPN ulanmazdan işleýär.
🎁 Promokod 1500 TMT: RICHONE (Hasaba alnanyňyzda… See the full description on the dataset page: https://huggingface.co/datasets/Andrey33312/1xbet-turkmenistan-zerkalo-tmcell.math_mesele_5mln_turkmenceturkmen_raw_textturkmen_dataset
