datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
risale-sohbet-turkish-2turkish-tts-combined-raw
Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından düzenlenmiştir. Orijinal veri seti afkfatih tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir.
🔗 Orijinal Kaynak: afkfatih/turkish-tts-combined-raw
🔗 Derleyen Platform: VeriPazarı
Türkçe TTS Birleşik Veri Seti (Turkish TTS Combined)
7 farklı açık kaynak Türkçe TTS (Metinden Sese) veri setinin birleşimidir.
~81.500 örnek |… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/turkish-tts-combined-raw.Turkish_TTS_Dataturkish-tts-combined-raw
Türkçe TTS Birleşik Veri Seti
7 farklı açık kaynak Türkçe TTS veri setinin birleşimi. ~81,500 örnek | 24kHz | SNAC uyumlu
Kaynaklar
Veri Seti
Örnek
Kaynak
Mazlum Kiper
9,643
omersaidd/tts_mazlum_kiper_tur
Ahmet Deniz
11,289
omersaidd/tts_ahmet_deniz_tur
Nisan Kumru
8,042
omersaidd/tts_nisan_kumru_tur
Derya TTS v2
42
afkfatih/derya-tts-v2
Derya Karma v3
255
afkfatih/derya-tts-karma-v3
Khan Academy
25,741
ysdede/khanacademy-turkish
Common… See the full description on the dataset page: https://huggingface.co/datasets/projectkaira/turkish-tts-combined-raw.turkish-audiobook-raw
Turkish Audiobook Speech Corpus (Raw)
Türkçe konuşma araştırmaları için derlenmiş, işlenmemiş uzun-form ses kayıtlarından
oluşan bir koleksiyon. Kayıtlar çeşitli kaynaklardan bir araya getirilmiştir ve
konuşmacı, kayıt ortamı, süre ve ses kalitesi bakımından geniş bir çeşitlilik gösterir.
İçerik
Uzun-form Türkçe konuşma kayıtları (m4a / mp3)
Kaynağa göre klasörlenmiş düz dizin yapısı
Transkript, hizalama veya segmentasyon içermez — ham hâldedir… See the full description on the dataset page: https://huggingface.co/datasets/serdarcaglar/turkish-audiobook-raw.turkish-tts-combined-raw
Türkçe TTS Birleşik Veri Seti
7 farklı açık kaynak Türkçe TTS veri setinin birleşimi. ~81,500 örnek | 24kHz | SNAC uyumlu
Kaynaklar
Veri Seti
Örnek
Kaynak
Mazlum Kiper
9,643
omersaidd/tts_mazlum_kiper_tur
Ahmet Deniz
11,289
omersaidd/tts_ahmet_deniz_tur
Nisan Kumru
8,042
omersaidd/tts_nisan_kumru_tur
Derya TTS v2
42
afkfatih/derya-tts-v2
Derya Karma v3
255
afkfatih/derya-tts-karma-v3
Khan Academy
25,741
ysdede/khanacademy-turkish
Common Voice 17
26… See the full description on the dataset page: https://huggingface.co/datasets/afkfatih/turkish-tts-combined-raw.turkish-tts-audiobooks
Turkish TTS Audiobooks
Turkish read-speech corpus for text-to-speech training, built from Turkish
audiobook and spoken-article recordings by an automatic pipeline: VAD
segmentation → technical QC → acoustic event tagging → DNSMOS → speaker
embedding/consistency → double-pass Whisper ASR → text policy → leakage-free
splitting. Audio is 16 kHz mono lossless FLAC embedded in the Parquet shards.
The pipeline that produced it — every stage, every threshold, the export and
audit… See the full description on the dataset page: https://huggingface.co/datasets/serdarcaglar/turkish-tts-audiobooks.turkishvoicedataset
Dataset Card for "turkishneuralvoice"
Dataset Overview
Dataset Name: Turkish Neural Voice
Description: This dataset contains Turkish audio samples generated using Microsoft Text to Speech services. The dataset includes audio files and their corresponding transcriptions.
Dataset Structure
Configs:
default
Data Files:
Split: train
Path: data/train-*
Dataset Info:
Features:
audio: Audio file
transcription: Corresponding text transcription
Splits:
train… See the full description on the dataset page: https://huggingface.co/datasets/erenfazlioglu/turkishvoicedataset.Turkish_Speech_Corpus
Turkish Speech Corpus (TSC)
This repository presents an open-source Turkish Speech Corpus, introduced in "Multilingual Speech Recognition for Turkic Languages". The corpus contains 218.2 hours of transcribed speech with 186,171 utterances and is the largest publicly available Turkish dataset of its kind at that time.
Paper: Multilingual Speech Recognition for Turkic Languages.
GitHub Repository: https://github.com/IS2AI/TurkicASR
Citation
@Article{info14020074… See the full description on the dataset page: https://huggingface.co/datasets/issai/Turkish_Speech_Corpus.turkish_male700h-tr-turkish-text-to-speechturkish-tts-kikiriturkish-tts-arena-v1turkish-tts-combined-raw
Türkçe TTS Birleşik Veri Seti
7 farklı açık kaynak Türkçe TTS veri setinin birleşimi. ~81,500 örnek | 24kHz | SNAC uyumlu
Kaynaklar
Veri Seti
Örnek
Kaynak
Mazlum Kiper
9,643
omersaidd/tts_mazlum_kiper_tur
Ahmet Deniz
11,289
omersaidd/tts_ahmet_deniz_tur
Nisan Kumru
8,042
omersaidd/tts_nisan_kumru_tur
Derya TTS v2
42
afkfatih/derya-tts-v2
Derya Karma v3
255
afkfatih/derya-tts-karma-v3
Khan Academy
25,741
ysdede/khanacademy-turkish
Common… See the full description on the dataset page: https://huggingface.co/datasets/Hm12cbbcbx/turkish-tts-combined-raw.khanacademy-turkish
Khan Academy Turkish Audio Dataset
This dataset contains 78 hours of audio extracted from the Khan Academy Turkish YouTube channel. The data has been segmented into short clips, each with an average duration of 10.5 seconds.
Accompanying this dataset, you will find a detailed video file tree that provides an overview of the source material.
Dataset Creation Process:The audio was extracted from the Khan Academy Turkish YouTube channel and then processed using several techniques to… See the full description on the dataset page: https://huggingface.co/datasets/ysdede/khanacademy-turkish.Turkish-Podcast-Merge-v1risale-sohbet-turkish
YouTube Transkripsiyon Veri Seti
Veri Yapısı
audio/: MP3 dosyaları
transcripts/: Metin transkripsiyonları
srt/: Altyazı dosyaları
metadata/: Video bilgileri
database.json: Tüm videoların indeksi
Güncelleme Tarihi
2025-03-21
TurkishVoiceDataset
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/cubukcum/TurkishVoiceDataset.Synthetic_Turkish_TTS_Data
Synthetic Turkish TTS Data
This dataset was created by generating synthetic Turkish text across multiple speech scenarios. The text was produced in the following domains: finance_master, cs_master, parcel_delivery, ecommerce, telecom, isp_support, technical_support, subscription, insurance, health_appointments, public_services, education_registration, and daily_speech.
These synthetic texts were then synthesized with a high-quality Turkish TTS model. The dataset is intended to be… See the full description on the dataset page: https://huggingface.co/datasets/Anilosan15/Synthetic_Turkish_TTS_Data.medv3-turkish-medical-asr
medv3 - Türkçe Sentetik Tıbbi Konuşma Korpusu
Türkçe tıbbi konuşma tanıma araştırmaları için hazırlanmış sentetik konuşma korpusudur.
Klinik cümleler Google Cloud Text-to-Speech Chirp 3 HD sesleriyle sentezlenmiştir.
Önemli uyarılar
Tüm kayıtlar sentetiktir (synthetic=true).
Gerçek hasta veya klinisyen sesi ve kişisel sağlık verisi içermez.
Tıbbi cihaz geliştirme onayı veya klinik doğrulama anlamına gelmez.
Klinik karar için değil, araştırma ve ASR… See the full description on the dataset page: https://huggingface.co/datasets/turkmedstt/medv3-turkish-medical-asr.khanacademy-turkish
Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından Türkçeye çevrilmiştir. Orijinal veri seti ysdede tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir.
🔗 Orijinal Kaynak: ysdede/khanacademy-turkish
🔗 Derleyen Platform: VeriPazarı
Khan Academy Türkçe Ses Veri Seti
Bu veri seti, Khan Academy Türkçe YouTube kanalından elde edilmiş 78 saatlik ses kaydını içermektedir. Veriler, her… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/khanacademy-turkish.turkishvoicedataset
Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından Türkçeye çevrilmiştir. Orijinal veri seti erenfazlioglu tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir.
🔗 Orijinal Kaynak: erenfazlioglu/turkishvoicedataset
🔗 Derleyen Platform: VeriPazarı
Türkçe Nöral Ses Veri Seti (Turkish Neural Voice)
Veri Setine Genel Bakış
Veri Seti Adı: Turkish Neural Voice (Türkçe… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/turkishvoicedataset.khanacademy-turkish-mathturkish-parliament-speechturkish-speech-datasetTurkish-Bentropi-tts-datasetturkish_male_10kturkish_female_10kTurkish-Podcast-8Turkish-Podcast-Merge
