datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
japanese-listening-voicevox-backupunified-kannada-asr-1.0
Dataset Card for "unified-kannada-asr-1.0"
More Information needed
hebrew_speech_kan
Dataset Card for Dataset Name
Dataset Summary
Hebrew Dataset for ASR
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
{'audio': {'path': '/root/.cache/huggingface/datasets/downloads/extracted/8ce7402f6482c6053251d7f3000eec88668c994beb48b7ca7352e77ef810a0b6/train/e429593fede945c185897e378a5839f4198.wav',
'array': array([-0.00265503, -0.0018158… See the full description on the dataset page: https://huggingface.co/datasets/imvladikon/hebrew_speech_kan.commonvoice_kana_onlyAgent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/kanepi-1977/Agent-Reasoning-WebSearch-260K.IndicTTS_Kannada
Kannada Indic TTS Dataset
This dataset is derived from the Indic TTS Database project, specifically using the Kannada monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development.
Dataset Details
Language: Kannada
Total Duration: ~7.35 hours (Male: 3.4 hours, Female: 3.95 hours)
Audio Format: WAV
Sampling Rate:… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Kannada.kana-sounds
Kana Sounds
147 short spoken clips, one for every hiragana and katakana character used by
Kana Trainer: the 46 seion, 20 dakuon,
5 handakuon, 33 yoon and 43 tokushon. They come from a single reader on
FUN Japanese Learning.
Dataset structure
audio/
seion/ 46 clips a.mp3, i.mp3, ka.mp3, ... n.mp3
dakuon/ 20 clips ga.mp3, za.mp3, ji.mp3, ... bo.mp3
handakuon/ 5 clips pa.mp3, pi.mp3, pu.mp3, pe.mp3, po.mp3
yoon/ 33 clips kya.mp3… See the full description on the dataset page: https://huggingface.co/datasets/arsalan-anwari/kana-sounds.syspin-kannada-ttskannada-speech-datasetiisc-mile-kannada-asr-corpuskannada_new_dataoriginal_data_kannada_ttsDarijaTTS-cleanvessel-detection-datasetKangaroo_20260115_20260413MOSS-TTS-ky-kk-bench
MOSS-TTS Kyrgyz/Kazakh Cross-Lingual Benchmark
Paired renderings of the same prompts by two text-to-speech models: MOSS-TTS v1.5
as released, and the same model with a Kyrgyz/Kazakh QLoRA adapter. Each row puts
the two side by side, so the effect of the fine-tune can be judged by ear rather than
from a metric.
The grid is deliberately cross-lingual: every reference voice is used with every
language, so an English speaker reads Kyrgyz, a Kazakh speaker reads Russian, and so on.… See the full description on the dataset page: https://huggingface.co/datasets/KaniTTS-research-team/MOSS-TTS-ky-kk-bench.dataset-from-restoreILRDF_Dict_Kanakanavu
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ILRDF_Dict_Kanakanavu
This is a noncanonical compatibility mirror. Use FormosanBank/ILRDF_Dicts for the complete canonical dataset and stable download contract.
This mirror… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ILRDF_Dict_Kanakanavu.kannada_new_data_v2emolia_filtered_v1
Emolia Filtered v1 (103,521 samples)
Subset of laion/Emolia processed through
the audio_filter pipeline.
All samples preserved (good + bad + uncertain), with filter results as additional columns.
Pipeline
Stage
Model
Purpose
1. Quality
Dual LogisticRegression (V1 SR<=24kHz / V3 SR>24kHz) on DSP metrics
Detect noise, clipping, robotic, bandwidth-limited audio
2. Speaker
Pyannote ONNX segmentation-3.0
Detect overlapping speakers
Speaker filter runs only on… See the full description on the dataset page: https://huggingface.co/datasets/KaniTTS-research-team/emolia_filtered_v1.central-kanuri-speech-datasetanguage:
kau
license: cc-by-nc-4.0
pretty_name: Central Kanuri Speech Dataset
task_categories:
automatic-speech-recognition
tags:
kanuri
central-kanuri
speech
audio
asr
low-resource-language
african-languages
speech-recognition
conversational-ai
Central Kanuri Speech Dataset
Overview
The Central Kanuri Speech Dataset is a community-contributed speech corpus developed by CIATECH Africa in collaboration with CLEAR Global through the TWB Voice initiative.
The dataset was developed to increase the… See the full description on the dataset page: https://huggingface.co/datasets/ciatech-frica/central-kanuri-speech-dataset.kanak30-ttsFineTune_Kannadakannada_datasetkannada-tts-annotatedIndicVoices-kannada-10000stt_synthetic_kn-IN_kannadaSPRING_INX_Kannada_R1kannada_new_data_v5kantipur-interview-data3
Nepali Speech Dataset (YouTube-sourced)
83 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split.
Splits
train: 83 segments
validation: 0 segments
test: 0 segments
Transcript columns — read this before training
Each segment carries three transcript variants. They are NOT interchangeable:
text_original — the YouTube caption text (if any) that overlapped this segment's… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/kantipur-interview-data3.
