datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arabic-audio-collection-mohamed-khairy
Mohamed Khairy Arabic Speech Dataset
Dataset Summary
The Mohamed Khairy Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 430 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-mohamed-khairy.khanacademy-turkish
Khan Academy Turkish Audio Dataset
This dataset contains 78 hours of audio extracted from the Khan Academy Turkish YouTube channel. The data has been segmented into short clips, each with an average duration of 10.5 seconds.
Accompanying this dataset, you will find a detailed video file tree that provides an overview of the source material.
Dataset Creation Process:The audio was extracted from the Khan Academy Turkish YouTube channel and then processed using several techniques to… See the full description on the dataset page: https://huggingface.co/datasets/ysdede/khanacademy-turkish.khanacademy-turkish
Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından Türkçeye çevrilmiştir. Orijinal veri seti ysdede tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir.
🔗 Orijinal Kaynak: ysdede/khanacademy-turkish
🔗 Derleyen Platform: VeriPazarı
Khan Academy Türkçe Ses Veri Seti
Bu veri seti, Khan Academy Türkçe YouTube kanalından elde edilmiş 78 saatlik ses kaydını içermektedir. Veriler, her… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/khanacademy-turkish.common-voice-urdu-processed
🎙️ Common Voice Urdu (Processed)
Ready-to-use Urdu speech dataset for fine-tuning ASR models
Mozilla Common Voice → Preprocessed → Whisper-Ready ✨
📊 Dataset at a Glance
Split
Samples
Use
🏋️ Train
7,339
Model training
🔧 Validation
5,046
Hyperparameter tuning
🧪 Test
5,091
Final evaluation
Total
17,476
💡 Audio is pre-resampled to 16kHz — plug directly into Whisper!
🚀 Quick Start
from datasets import load_dataset
#… See the full description on the dataset page: https://huggingface.co/datasets/khawajaaliarshad/common-voice-urdu-processed.arabic-speech-SADA22-Khaliji
Dataset Card for SADA (Saudi Audio Dataset for Arabic)
⚠️ Caution
This is only the portion of the SADA dataset where the speaker dialect is Khaliji. To access full dataset, you should check this link.
Dataset Summary
The SADA dataset (Saudi Audio Dataset for Arabic) is a large-scale Arabic speech corpus designed to support the development of high-quality artificial intelligence models for Arabic speech processing. It contains over 667 hours of… See the full description on the dataset page: https://huggingface.co/datasets/badrex/arabic-speech-SADA22-Khaliji.bhasaflow-khasi-english-parallel-sample-v1
BhasaFlow Khasi-English Parallel Sample v1
A professionally curated, gold-standard parallel speech and text corpus for the Khasi language.
Published by Medharvix Systems Private Limited
Part of the BhasaFlow Low-Resource Language Technology Initiative
Overview
This repository contains a public sample preview of the BhasaFlow Khasi-English Parallel Corpus, a structured speech and text dataset developed by Medharvix Systems Private Limited. The dataset pairs… See the full description on the dataset page: https://huggingface.co/datasets/MEDHARVIX-SYSTEMS/bhasaflow-khasi-english-parallel-sample-v1.bhasaflow-khasi-english-parallel-sample-v1
BhasaFlow Khasi-English Parallel Sample v1
A professionally curated, gold-standard parallel speech and text corpus for the Khasi language.
Published by Medharvix Systems Private Limited
Part of the BhasaFlow Low-Resource Language Technology Initiative
Overview
This repository contains a public sample preview of the BhasaFlow Khasi-English Parallel Corpus, a structured speech and text dataset developed by Medharvix Systems Private Limited. The dataset pairs… See the full description on the dataset page: https://huggingface.co/datasets/1infinity0/bhasaflow-khasi-english-parallel-sample-v1.dhivehi-khadheeja-speechDhivehi Khadheeja Speech is a single speaker Dhivehi speech dataset created by [Javaabu Pvt. Ltd.](https://javaabu.com).
The dataset contains around 20 hrs of text read by professional Maldivian narrator Khadheeja Faaz.
The text used for the recordings were text scrapped from various Maldivian news websites.Khasi-OmniVoice-TTS-Data
Khasi Omni Voice TTS Dataset
The Khasi Omni Voice dataset is a comprehensive, high-quality audio collection designed specifically for Text-to-Speech (TTS) research and model training in the Khasi language. It features nearly 50 hours of speech data targeting realistic, modern Khasi speech patterns, including natural code-switching.
Key Statistics
Total Duration: 49 hours, 53 minutes, 19.98 seconds
Total Samples: 18,874 distinct audio utterances
Language: Khasi… See the full description on the dataset page: https://huggingface.co/datasets/toiar/Khasi-OmniVoice-TTS-Data.Khasi_ASR_Dataset
Khasi ASR Dataset
The Khasi ASR Dataset is a large-scale speech recognition dataset for the Khasi language, an Indigenous language spoken primarily in Meghalaya, India. The dataset contains paired audio recordings and transcriptions designed for training and evaluating Automatic Speech Recognition (ASR) systems.
This dataset consists of 73,900 audio-transcription pairs with a total duration of approximately 101 hours, 19 minutes, and 54.36 seconds of speech data.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/toiar/Khasi_ASR_Dataset.
