CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mamed0v /TurkmenSpeech Turkmen Speech Dataset (ASR) This dataset contains 251 hours of Turkmen speech audio with transcriptions, intended for training and evaluating Automatic Speech Recognition (ASR) models. It is one of the largest publicly available Turkmen speech datasets. Dataset Overview Property Value Total clips 119,847 Total duration 251.86 hours Sampling rate 16,000 Hz Language Turkmen (tk) Split train Each item includes: audio: waveform + sampling… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/TurkmenSpeech.audioautomatic-speech-recognition100K<n<1M7 likes1.4k downloads11mo agoHugging Face02turiabu /Sagalee Sagalee – Automatic Speech Recognition Dataset for Afaan Oromoo Dataset Description Sagalee is Speech Recognition Dataset for Oromo language Presented in the paper: Sagalee: an Open Source Automatic Speech Recognition Dataset for Oromo Language at ICASSP 2025 Training Code: turinaaf/Sagalee Arxiv: https://arxiv.org/abs/2502.00421 The dataset is released under Attribution–NonCommercial 4.0 International (CC BY-NC 4.0) It contains read-speech recordings from native… See the full description on the dataset page: https://huggingface.co/datasets/turiabu/Sagalee.audioautomatic-speech-recognition10K<n<100K2 likes621 downloads7mo agoHugging Face03Anilosan15 /Turkish_TTS_Dataaudiotext-to-speech10K<n<100K21 likes448 downloads7mo agoHugging Face04serdarcaglar /turkish-audiobook-rawgated Turkish Audiobook Speech Corpus (Raw) Türkçe konuşma araştırmaları için derlenmiş, işlenmemiş uzun-form ses kayıtlarından oluşan bir koleksiyon. Kayıtlar çeşitli kaynaklardan bir araya getirilmiştir ve konuşmacı, kayıt ortamı, süre ve ses kalitesi bakımından geniş bir çeşitlilik gösterir. İçerik Uzun-form Türkçe konuşma kayıtları (m4a / mp3) Kaynağa göre klasörlenmiş düz dizin yapısı Transkript, hizalama veya segmentasyon içermez — ham hâldedir… See the full description on the dataset page: https://huggingface.co/datasets/serdarcaglar/turkish-audiobook-raw.audiotext-to-speechn<1K0 likes396 downloads27d agoHugging Face05serdarcaglar /turkish-tts-audiobooksgated Turkish TTS Audiobooks Turkish read-speech corpus for text-to-speech training, built from Turkish audiobook and spoken-article recordings by an automatic pipeline: VAD segmentation → technical QC → acoustic event tagging → DNSMOS → speaker embedding/consistency → double-pass Whisper ASR → text policy → leakage-free splitting. Audio is 16 kHz mono lossless FLAC embedded in the Parquet shards. The pipeline that produced it — every stage, every threshold, the export and audit… See the full description on the dataset page: https://huggingface.co/datasets/serdarcaglar/turkish-tts-audiobooks.audiotext-to-speech100K<n<1M9 likes380 downloads1mo agoHugging Face06issai /Turkish_Speech_Corpus Turkish Speech Corpus (TSC) This repository presents an open-source Turkish Speech Corpus, introduced in "Multilingual Speech Recognition for Turkic Languages". The corpus contains 218.2 hours of transcribed speech with 186,171 utterances and is the largest publicly available Turkish dataset of its kind at that time. Paper: Multilingual Speech Recognition for Turkic Languages. GitHub Repository: https://github.com/IS2AI/TurkicASR Citation @Article{info14020074… See the full description on the dataset page: https://huggingface.co/datasets/issai/Turkish_Speech_Corpus.audioautomatic-speech-recognition13 likes295 downloads2y agoHugging Face07futureDoctor /turkic_tts_dataset Turkic TTS Dataset A multilingual TTS corpus covering Turkic languages. Languages Subset Source Speakers azerbaijani BHOSAI/Azerbaijani_News_TTS 1 (female) bashkir AigizK/bashkort_tts_dataset 8 (7F + 1M, ElevenLabs cloned) Schema Column Type Description audio Audio Speech sample text string Transcription source_link string Original dataset URL speaker_idstring Speaker identifier (and style if applicable) gender string… See the full description on the dataset page: https://huggingface.co/datasets/futureDoctor/turkic_tts_dataset.audiotext-to-speech10K<n<100K0 likes281 downloads4mo agoHugging Face08Appenlimited /700h-tr-turkish-text-to-speechaudioautomatic-speech-recognition1K<n<10K17 likes207 downloads1y agoHugging Face09ysdede /khanacademy-turkish Khan Academy Turkish Audio Dataset This dataset contains 78 hours of audio extracted from the Khan Academy Turkish YouTube channel. The data has been segmented into short clips, each with an average duration of 10.5 seconds. Accompanying this dataset, you will find a detailed video file tree that provides an overview of the source material. Dataset Creation Process:The audio was extracted from the Khan Academy Turkish YouTube channel and then processed using several techniques to… See the full description on the dataset page: https://huggingface.co/datasets/ysdede/khanacademy-turkish.audioautomatic-speech-recognition10K<n<100K36 likes152 downloads2y agoHugging Face10cubukcum /TurkishVoiceDataset Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/cubukcum/TurkishVoiceDataset.audiotext-to-audio100K<n<1M4 likes119 downloads1y agoHugging Face11omersaidd /tts_ahmet_deniz_tur Merhaba Arkadaşlar 🚀 Türçe TTS üzerinde oluşturduğum veri seti bu şekildedir. Amacım TTS ( Text-To-Speech) üzerine açık kaynak modellerin Türkçe performansını iyileştirmek ve daha kaliteli çıktılar vermesini sağlamaktır. Bunun için çeşitli kaynaklardan elde ettiğim videoları doğru formata getirip eğitime hazır bir veri seti oluşturdum, Bu veri setinin oluşturmakta ki amacım ticari bir gaye değil, araştırma alanında çalışan kişilere fayda… See the full description on the dataset page: https://huggingface.co/datasets/omersaidd/tts_ahmet_deniz_tur.audiotext-to-speech10K<n<100K4 likes111 downloads9mo agoHugging Face12Anilosan15 /Synthetic_Turkish_TTS_Data Synthetic Turkish TTS Data This dataset was created by generating synthetic Turkish text across multiple speech scenarios. The text was produced in the following domains: finance_master, cs_master, parcel_delivery, ecommerce, telecom, isp_support, technical_support, subscription, insurance, health_appointments, public_services, education_registration, and daily_speech. These synthetic texts were then synthesized with a high-quality Turkish TTS model. The dataset is intended to be… See the full description on the dataset page: https://huggingface.co/datasets/Anilosan15/Synthetic_Turkish_TTS_Data.audiotext-to-speech10K<n<100K6 likes109 downloads5mo agoHugging Face13turkmedstt /medv3-turkish-medical-asr medv3 - Türkçe Sentetik Tıbbi Konuşma Korpusu Türkçe tıbbi konuşma tanıma araştırmaları için hazırlanmış sentetik konuşma korpusudur. Klinik cümleler Google Cloud Text-to-Speech Chirp 3 HD sesleriyle sentezlenmiştir. Önemli uyarılar Tüm kayıtlar sentetiktir (synthetic=true). Gerçek hasta veya klinisyen sesi ve kişisel sağlık verisi içermez. Tıbbi cihaz geliştirme onayı veya klinik doğrulama anlamına gelmez. Klinik karar için değil, araştırma ve ASR… See the full description on the dataset page: https://huggingface.co/datasets/turkmedstt/medv3-turkish-medical-asr.audioautomatic-speech-recognition1K<n<10K4 likes100 downloads3mo agoHugging Face14Taklaxbr /khanacademy-turkish Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından Türkçeye çevrilmiştir. Orijinal veri seti ysdede tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir. 🔗 Orijinal Kaynak: ysdede/khanacademy-turkish 🔗 Derleyen Platform: VeriPazarı Khan Academy Türkçe Ses Veri Seti Bu veri seti, Khan Academy Türkçe YouTube kanalından elde edilmiş 78 saatlik ses kaydını içermektedir. Veriler, her… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/khanacademy-turkish.audioautomatic-speech-recognition10K<n<100K0 likes98 downloads4mo agoHugging Face15alimetin /turkish-parliament-speechaudioautomatic-speech-recognition1K<n<10K1 likes73 downloads9mo agoHugging Face16Taklaxbr /Synthetic_Turkish_TTS_Data Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından Türkçeye çevrilmiştir. Orijinal veri seti Anilosan15 tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir. 🔗 Orijinal Kaynak: Anilosan15/Synthetic_Turkish_TTS_Data 🔗 Derleyen Platform: VeriPazarı Sentetik Türkçe TTS Veri Seti (Synthetic Turkish TTS Data) Bu veri seti, çoklu konuşma senaryoları üzerinden sentetik Türkçe metinler… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/Synthetic_Turkish_TTS_Data.audiotext-to-speech10K<n<100K1 likes45 downloads4mo agoHugging Face17speedykom-group /turkana-speech-dataset Turkana Speech Dataset Speech dataset for Turkana (tuv) — Eastern Nilotic language, ~1M speakers, Kenya. Property Value Format WAV, 16 kHz, mono / UTF-8 transcripts Clips 5,151 segments Splits Train: 3,090 (60%) · Validation: 1,030 (20%) · Test: 1,031 (20%) — seed 42 Source GRN Bible narratives (Global Recordings Network, LLL series 1–8), segmented via silence detection Transcription Auto-generated via facebook/mms-1b-all (Teso adapter)… See the full description on the dataset page: https://huggingface.co/datasets/speedykom-group/turkana-speech-dataset.audiotext-to-speech1K<n<10K1 likes28 downloads1mo agoHugging Face18Speech-data /Turkish-Speech-Dataset 🎧 Turkish Speech Dataset The Turkish Speech Dataset is a high-quality speech audio dataset designed to power modern AI and machine learning solutions with diverse and structured audio data. It includes 123 hours of voice recordings distributed across 802 files, provided in MP3 and WAV formats, with a total size of 105 MB. This carefully curated audio dataset delivers balanced and representative voice data, with 46% female and 54% male speakers, and an age range spanning from 18 to… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Turkish-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes25 downloads6mo agoHugging Face19Thomcles /YodaLingua-Turkishgated YodaLingua-Turkish YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Turkish portion of the multilingual YodaLingua collection. 🧾 Dataset Overview Property Value Total clips 56,422 audio–transcription pairs Total duration 163 hours Speakers 2,255 distinct speakers Audio format MP3 • mono • 24 kHz • 16-bit… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Turkish.audiotext-to-speech10K<n<100K1 likes21 downloads5mo agoHugging Face20Taklaxbr /Turkish_Speech_Corpus Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından Türkçeye çevrilmiştir. Orijinal veri seti issai tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir. 🔗 Orijinal Kaynak: issai/Turkish_Speech_Corpus 🔗 Derleyen Platform: VeriPazarı Turkish Speech Corpus (TSC) Bu depo, "Multilingual Speech Recognition for Turkic Languages" (Türkî Diller İçin Çok Dilli Konuşma Tanıma) makalesinde… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/Turkish_Speech_Corpus.audioautomatic-speech-recognition0 likes16 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.