datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TurkmenSpeech
Turkmen Speech Dataset (ASR)
This dataset contains 251 hours of Turkmen speech audio with transcriptions, intended for training and evaluating Automatic Speech Recognition (ASR) models.
It is one of the largest publicly available Turkmen speech datasets.
Dataset Overview
Property
Value
Total clips
119,847
Total duration
251.86 hours
Sampling rate
16,000 Hz
Language
Turkmen (tk)
Split
train
Each item includes:
audio: waveform + sampling… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/TurkmenSpeech.Easy-Turn-Trainset
Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems
Guojian Li1, Chengyou Wang1, Hongfei Xue1,
Shuiyuan Wang1, Dehui Gao1, Zihan Zhang2,
Yuke Lin2, Wenjie Li2, Longshuai Xiao2,
Zhonghua Fu1,╀, Lei Xie1,╀
1 Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University
2 Huawei Technologies, China
🎤 Demo Page
🤖 Easy Turn Model
📑 Paper
🌐 Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/Easy-Turn-Trainset.Sagalee
Sagalee – Automatic Speech Recognition Dataset for Afaan Oromoo
Dataset Description
Sagalee is Speech Recognition Dataset for Oromo language Presented in the paper:
Sagalee: an Open Source Automatic Speech Recognition Dataset for Oromo Language at ICASSP 2025
Training Code: turinaaf/Sagalee
Arxiv: https://arxiv.org/abs/2502.00421
The dataset is released under Attribution–NonCommercial 4.0 International (CC BY-NC 4.0) It contains read-speech recordings from native… See the full description on the dataset page: https://huggingface.co/datasets/turiabu/Sagalee.Real-TurnTurk
Real-TurnTurk
English: Real-TurnTurk is a multimodal, two-channel Turkish dyadic conversation dataset built to improve turn-taking prediction in voice-based dialogue systems. Unlike Syn-TurnTurk, the other dataset we built, every conversation here is a real, unscripted exchange between two people, recorded over video calls. Each participant was captured on a separate audio channel, so speaker attribution is exact and requires no diarization model. Alongside the audio, the… See the full description on the dataset page: https://huggingface.co/datasets/tugrulbayrak/Real-TurnTurk.Turkish_TTS_Dataturkish-tts-audiobooks
Turkish TTS Audiobooks
Turkish read-speech corpus for text-to-speech training, built from Turkish
audiobook and spoken-article recordings by an automatic pipeline: VAD
segmentation → technical QC → acoustic event tagging → DNSMOS → speaker
embedding/consistency → double-pass Whisper ASR → text policy → leakage-free
splitting. Audio is 16 kHz mono lossless FLAC embedded in the Parquet shards.
The pipeline that produced it — every stage, every threshold, the export and
audit… See the full description on the dataset page: https://huggingface.co/datasets/serdarcaglar/turkish-tts-audiobooks.turkic_tts_dataset
Turkic TTS Dataset
A multilingual TTS corpus covering Turkic languages.
Languages
Subset
Source
Speakers
azerbaijani
BHOSAI/Azerbaijani_News_TTS
1 (female)
bashkir
AigizK/bashkort_tts_dataset
8 (7F + 1M, ElevenLabs cloned)
Schema
Column
Type
Description
audio
Audio
Speech sample
text
string
Transcription
source_link
string
Original dataset URL
speaker_idstring
Speaker identifier (and style if applicable)
gender
string… See the full description on the dataset page: https://huggingface.co/datasets/futureDoctor/turkic_tts_dataset.700h-tr-turkish-text-to-speechkhanacademy-turkish
Khan Academy Turkish Audio Dataset
This dataset contains 78 hours of audio extracted from the Khan Academy Turkish YouTube channel. The data has been segmented into short clips, each with an average duration of 10.5 seconds.
Accompanying this dataset, you will find a detailed video file tree that provides an overview of the source material.
Dataset Creation Process:The audio was extracted from the Khan Academy Turkish YouTube channel and then processed using several techniques to… See the full description on the dataset page: https://huggingface.co/datasets/ysdede/khanacademy-turkish.TurnMaster
Download
# Install huggingface_hub if needed
pip install huggingface_hub
# Download dataset
hf download turnmaster/TurnMaster --repo-type dataset --local-dir ./TurnMaster
cd TurnMaster
Installation
pip install -r requirements.txt
Project structure
TurnMaster/
├── audio/
├── metadata/
├── processing/
│ ├── augment_taskmaster.py
│ ├── extract_shards.py
│ ├── curate_dataset.py
├── requirements.txt
└── README.md
Extract shards… See the full description on the dataset page: https://huggingface.co/datasets/turnmaster/TurnMaster.TurkishVoiceDataset
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/cubukcum/TurkishVoiceDataset.tamil-turns
tamil-turns
4,774 conversational turns from 115 real Tamil telephone calls —
one speaker's continuous hold of the floor, with their own mid-turn pauses kept
separate from the silence that ends the turn.
Companion dataset.
santhosh-005/tamil-eot
is the same source calls cut as 8 s windows with a binary complete /
incomplete label — use that one to train an end-of-turn classifier.
This dataset keeps whole turns with their silence structure — use it
for turn-taking, mid-turn… See the full description on the dataset page: https://huggingface.co/datasets/santhosh-005/tamil-turns.tts_ahmet_deniz_tur
Merhaba Arkadaşlar 🚀
Türçe TTS üzerinde oluşturduğum veri seti bu şekildedir.
Amacım TTS ( Text-To-Speech) üzerine açık kaynak modellerin Türkçe performansını iyileştirmek ve daha kaliteli çıktılar vermesini sağlamaktır. Bunun için çeşitli kaynaklardan elde ettiğim videoları doğru formata getirip eğitime hazır bir veri seti oluşturdum,
Bu veri setinin oluşturmakta ki amacım ticari bir gaye değil, araştırma alanında çalışan kişilere fayda… See the full description on the dataset page: https://huggingface.co/datasets/omersaidd/tts_ahmet_deniz_tur.Synthetic_Turkish_TTS_Data
Synthetic Turkish TTS Data
This dataset was created by generating synthetic Turkish text across multiple speech scenarios. The text was produced in the following domains: finance_master, cs_master, parcel_delivery, ecommerce, telecom, isp_support, technical_support, subscription, insurance, health_appointments, public_services, education_registration, and daily_speech.
These synthetic texts were then synthesized with a high-quality Turkish TTS model. The dataset is intended to be… See the full description on the dataset page: https://huggingface.co/datasets/Anilosan15/Synthetic_Turkish_TTS_Data.medv3-turkish-medical-asr
medv3 - Türkçe Sentetik Tıbbi Konuşma Korpusu
Türkçe tıbbi konuşma tanıma araştırmaları için hazırlanmış sentetik konuşma korpusudur.
Klinik cümleler Google Cloud Text-to-Speech Chirp 3 HD sesleriyle sentezlenmiştir.
Önemli uyarılar
Tüm kayıtlar sentetiktir (synthetic=true).
Gerçek hasta veya klinisyen sesi ve kişisel sağlık verisi içermez.
Tıbbi cihaz geliştirme onayı veya klinik doğrulama anlamına gelmez.
Klinik karar için değil, araştırma ve ASR… See the full description on the dataset page: https://huggingface.co/datasets/turkmedstt/medv3-turkish-medical-asr.khanacademy-turkish
Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından Türkçeye çevrilmiştir. Orijinal veri seti ysdede tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir.
🔗 Orijinal Kaynak: ysdede/khanacademy-turkish
🔗 Derleyen Platform: VeriPazarı
Khan Academy Türkçe Ses Veri Seti
Bu veri seti, Khan Academy Türkçe YouTube kanalından elde edilmiş 78 saatlik ses kaydını içermektedir. Veriler, her… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/khanacademy-turkish.paralingua_ru
Russian Paralinguistic Annotation Dataset
Датасет паралингвистической разметки спикеров из трёх русскоязычных корпусов:
biggest_ru_book,
DeepSpeech и Golos.
Что размечалось
Каждое аудио размечалось вручную по следующим характеристикам:
Поле
Описание
Пример значений
gender
Пол спикера
мужской, женский
age_group
Возрастная группа
молодой, взрослый, пожилой
voice_pitch
Высота голоса
низкий, средний, высокий
loudness
Громкость
тихий, нормальный… See the full description on the dataset page: https://huggingface.co/datasets/turnipseason/paralingua_ru.candor-turntaking-annotations
CANDOR - Turn-Taking Annotations
Speech transcription and turn-taking annotation dataset built from the CANDOR corpus using NVIDIA Canary-Qwen2.5B ASR.
Dataset Description
This dataset contains 172,591 transcribed speech segments from the CANDOR conversational speech corpus (1,656 conversations). Each segment is a per-speaker utterance with Canary ASR transcript, designed for turn-taking prediction research.
Source
Audio corpus: CANDOR (English conversational… See the full description on the dataset page: https://huggingface.co/datasets/hiraki/candor-turntaking-annotations.turkish-parliament-speechSynthetic_Turkish_TTS_Data
Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından Türkçeye çevrilmiştir. Orijinal veri seti Anilosan15 tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir.
🔗 Orijinal Kaynak: Anilosan15/Synthetic_Turkish_TTS_Data
🔗 Derleyen Platform: VeriPazarı
Sentetik Türkçe TTS Veri Seti (Synthetic Turkish TTS Data)
Bu veri seti, çoklu konuşma senaryoları üzerinden sentetik Türkçe metinler… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/Synthetic_Turkish_TTS_Data.Easy-Turn-Trainset
Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems
Guojian Li1, Chengyou Wang1, Hongfei Xue1,
Shuiyuan Wang1, Dehui Gao1, Zihan Zhang2,
Yuke Lin2, Wenjie Li2, Longshuai Xiao2,
Zhonghua Fu1,╀, Lei Xie1,╀
1 Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University
2 Huawei Technologies, China
🎤 Demo Page
🤖 Easy Turn Model
📑 Paper
🌐 Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/0x3/Easy-Turn-Trainset.turkana-speech-dataset
Turkana Speech Dataset
Speech dataset for Turkana (tuv) — Eastern Nilotic language, ~1M speakers, Kenya.
Property
Value
Format
WAV, 16 kHz, mono / UTF-8 transcripts
Clips
5,151 segments
Splits
Train: 3,090 (60%) · Validation: 1,030 (20%) · Test: 1,031 (20%) — seed 42
Source
GRN Bible narratives (Global Recordings Network, LLL series 1–8), segmented via silence detection
Transcription
Auto-generated via facebook/mms-1b-all (Teso adapter)… See the full description on the dataset page: https://huggingface.co/datasets/speedykom-group/turkana-speech-dataset.Turkish-Speech-Dataset
🎧 Turkish Speech Dataset
The Turkish Speech Dataset is a high-quality speech audio dataset designed to power modern AI and machine learning solutions with diverse and structured audio data. It includes 123 hours of voice recordings distributed across 802 files, provided in MP3 and WAV formats, with a total size of 105 MB. This carefully curated audio dataset delivers balanced and representative voice data, with 46% female and 54% male speakers, and an age range spanning from 18 to… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Turkish-Speech-Dataset.YodaLingua-Turkish
YodaLingua-Turkish
YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Turkish portion of the multilingual YodaLingua collection.
🧾 Dataset Overview
Property
Value
Total clips
56,422 audio–transcription pairs
Total duration
163 hours
Speakers
2,255 distinct speakers
Audio format
MP3 • mono • 24 kHz • 16-bit… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Turkish.
