datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Urdu-ONYX-WAV-kanade-Annotated
Urdu-ONYX-WAV-real-Annotated
Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features.
Dataset Statistics
Total Samples: 26,217
Total Duration: 42.77 hours
Average Duration: 5.87 seconds
Duration Range: 0.65s - 122.23s
Average Phonemes: 18.5 per sample
Average Kanade Tokens: 151.1 per sample
Global Embedding Dimension: 128
New Columns
This dataset adds the following columns:
duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.hebrew_speech_kan
Dataset Card for Dataset Name
Dataset Summary
Hebrew Dataset for ASR
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
{'audio': {'path': '/root/.cache/huggingface/datasets/downloads/extracted/8ce7402f6482c6053251d7f3000eec88668c994beb48b7ca7352e77ef810a0b6/train/e429593fede945c185897e378a5839f4198.wav',
'array': array([-0.00265503, -0.0018158… See the full description on the dataset page: https://huggingface.co/datasets/imvladikon/hebrew_speech_kan.kana-sounds
Kana Sounds
147 short spoken clips, one for every hiragana and katakana character used by
Kana Trainer: the 46 seion, 20 dakuon,
5 handakuon, 33 yoon and 43 tokushon. They come from a single reader on
FUN Japanese Learning.
Dataset structure
audio/
seion/ 46 clips a.mp3, i.mp3, ka.mp3, ... n.mp3
dakuon/ 20 clips ga.mp3, za.mp3, ji.mp3, ... bo.mp3
handakuon/ 5 clips pa.mp3, pi.mp3, pu.mp3, pe.mp3, po.mp3
yoon/ 33 clips kya.mp3… See the full description on the dataset page: https://huggingface.co/datasets/arsalan-anwari/kana-sounds.kannada_asr_corpusThe corpus contains roughly 360 hours of audio and transcripts in Kannada language. The transcripts have beed de-duplicated using exact match deduplication.kantipur-interview-data3
Nepali Speech Dataset (YouTube-sourced)
83 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split.
Splits
train: 83 segments
validation: 0 segments
test: 0 segments
Transcript columns — read this before training
Each segment carries three transcript variants. They are NOT interchangeable:
text_original — the YouTube caption text (if any) that overlapped this segment's… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/kantipur-interview-data3.kantipur-interview-data1
Nepali Speech Dataset (YouTube-sourced)
204 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split.
Splits
train: 204 segments
validation: 0 segments
test: 0 segments
Transcript columns — read this before training
Each segment carries three transcript variants. They are NOT interchangeable:
text_original — the YouTube caption text (if any) that overlapped this… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/kantipur-interview-data1.kantipur-interview-data2
Nepali Speech Dataset (YouTube-sourced)
157 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split.
Splits
train: 157 segments
validation: 0 segments
test: 0 segments
Transcript columns — read this before training
Each segment carries three transcript variants. They are NOT interchangeable:
text_original — the YouTube caption text (if any) that overlapped this… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/kantipur-interview-data2.Kannada_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 16,549 hours of processed Kannada (KN) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Kannada_Call_Center_Audio_Dataset_Dual_Channel.conversational_kannada_stt
Conversational Kannada STT
This dataset contains corrected transcriptions of conversational Kannada speech,prepared specifically for fine-tuning Whisper models on conversational and dialectal Kannada.
Unlike many ASR datasets, this release provides pre-computed Whisper input features (log-Mel spectrograms)so you can train/fine-tune Whisper models without raw audio processing.
Dataset Creation
Source Audio: Publicly available YouTube videos in Kannada.
Initial… See the full description on the dataset page: https://huggingface.co/datasets/ShimogaAIteam/conversational_kannada_stt.kantipur-news-data
Nepali Speech Dataset (YouTube-sourced)
74 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split.
Splits
train: 74 segments
validation: 0 segments
test: 0 segments
Transcript columns — read this before training
Each segment carries three transcript variants. They are NOT interchangeable:
text_original — the YouTube caption text (if any) that overlapped this segment's… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/kantipur-news-data.Kannada-Call-Center-Audio-Dataset-Single-ChannelDataset Description:
This dataset is a large-scale collection of 16,549 hours of processed Kannada (KN) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural speaking behaviour commonly observed in… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Kannada-Call-Center-Audio-Dataset-Single-Channel.Kannada-Speech-Dataset
🎧 Kannada Speech Dataset
The Kannada Speech Dataset is a high-quality speech audio dataset designed to deliver structured and reliable audio data for AI and machine learning workflows. It includes 90 hours of audio data across 651 files, available in MP3 and WAV formats, with a total size of 220 MB. This well-organized audio dataset provides balanced and representative voice data, with 48% female and 52% male speakers, and an age range spanning from 18 to 50+ years. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Kannada-Speech-Dataset.Fleurs-KnThis is a filtered version of the Fleurs dataset only containing samples of Kannada language.
The dataset contains total of 2283 training, 368 validation and 838 test samples.
Data Sample:
{'id': 1053,
'num_samples': 226560,
'path': '/home/ravi.naik/.cache/huggingface/datasets/downloads/extracted/e7c8b501d4e6892673b6dc291d42de48e7987b0d2aa6471066a671f686224ed1/10000267636955490843.wav',
'audio': {'path': 'train/10000267636955490843.wav',
'array': array([ 0. , 0.… See the full description on the dataset page: https://huggingface.co/datasets/Kannada-LLM-Labs/Fleurs-Kn.Kannada_Podcast_Audio_Dataset_Dual_Channel
Dataset Description
This dataset is a large-scale collection of 3,970 hours of processed Kannada dual-channel podcast audio recordings, containing 57,569 hours of processed podcast audio recordings across 12 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It captures real-world podcast conversations across diverse topics and formats. The dataset is organized in a dual-channel format, where corresponding speaker… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Kannada_Podcast_Audio_Dataset_Dual_Channel.Urdu-ONYX-WAV-kanade-V2
Urdu-ONYX-WAV-kanade-Annotated-V2
Version 2.0 - Artifact-Free Edition 🎉
Overview
This is an improved version of the Urdu-ONYX-WAV dataset, tokenized with the Kanade neural codec and optimized for artifact-free audio decoding. This dataset contains 143,627 samples of high-quality Urdu speech with comprehensive linguistic and acoustic annotations, totaling ~244 hours (~10 days) of continuous audio.
Key Features
🎯 Large-Scale: 143K+ samples, 244+ hours of… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-V2.
