datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ISSAI_KSC_335RS_v_1_1
Dataset Card for "ISSAI_KSC_335RS_v_1_1"
Kazakh Speech Corpus (KSC)
Identifier: SLR102
Summary: A crowdsourced open-source Kazakh speech corpus developed by ISSAI (330 hours)
Category: Speech
License: Attribution 4.0 International (CC BY 4.0)
Downloads (use a mirror closer to you):
ISSAI_KSC_335RS_v1.1_flac.tar.gz [19G] (speech, transcripts and metadata ) Mirrors: [US] [EU] [CN]
About this resource:
A crowdsourced open-source speech corpus for the Kazakh language. The KSC… See the full description on the dataset page: https://huggingface.co/datasets/Shirali/ISSAI_KSC_335RS_v_1_1.TTSDistil-Phonology
Overview
This repository contains a speech dataset developed for Text-to-Speech (TTS) research and model training.
The dataset is part of an ongoing research project focused on building high-quality speech corpora for modern neural TTS systems. It is actively maintained, with continuous improvements in data quality, transcription consistency, metadata, and organization.
This repository is intended to host research data used throughout the development process and is not intended… See the full description on the dataset page: https://huggingface.co/datasets/ShiniChien/TTSDistil-Phonology.TTSDistil2
Overview
This repository contains a speech dataset developed for Text-to-Speech (TTS) research and model training.
The dataset is part of an ongoing research project focused on building high-quality speech corpora for modern neural TTS systems. It is actively maintained, with continuous improvements in data quality, transcription consistency, metadata, and organization.
This repository is intended to host research data used throughout the development process and is not intended… See the full description on the dataset page: https://huggingface.co/datasets/ShiniChien/TTSDistil2.ePark_tu_hua_gu_shi_pian_picture_story
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_tu_hua_gu_shi_pian_picture_story
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.
This is a… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_tu_hua_gu_shi_pian_picture_story.ATC_combined
Dataset Card for UWB-ATCC corpus
Dataset Summary
The UWB-ATCC Corpus is provided provided by University of West Bohemia, Department of Cybernetics. The corpus contains recordings of communication between air traffic controllers and pilots. The speech is manually transcribed and labeled with the information about the speaker (pilot/controller, not the full identity of the person). The corpus is currently small (20 hours) but we plan to search for additional data next year.… See the full description on the dataset page: https://huggingface.co/datasets/Shiry/ATC_combined.Shinekhen-BuryatAudio collected by Yamakoshi (Tokyo University of Foreign Studies), originally uploaded here (CC BY-SA 4.0).
start_time and end_time are from the original audio clips; the audio uploaded here are already converted into per-sentence audio clips.
Used in [paper] [GitHub]
nirantar
Nirantar
Nirantar speech dataset (22 Indian languages) in Hugging Face format. Source: AI4Bharat/Nirantar.
Downloading language-wise subsets
Each language is a separate configuration (subset), so you can load only one language (like ai4bharat/Rasa):
from datasets import load_dataset, get_dataset_config_names
# Single language (only that language's parquet is downloaded)
ds = load_dataset("adjaysagar/nirantar", "hi", trust_remote_code=True) # Hindi
ds["train"] #… See the full description on the dataset page: https://huggingface.co/datasets/shiprocket-ai/nirantar.VoiceConvDesign
Dataset columns
Column
Type
Description
conversation_id
string
Unique identifier for the conversation
turn_index
int32
Zero-based turn position within the conversation
agent
string
Agent name
voice
string
TTS voice used for this turn
prompt
string
Full system instruction sent to Gemini Live for this speaker
transcript
string
Text spoken in this turn
duration_s
float32
Approximate audio duration in seconds
audio
audio
WAV audio (24 kHz mono 16-bit)OmniEdu
Overview
This dataset is being developed for training and evaluating omni-modal speech models, with a primary focus on audio-to-audio tasks. The repository is part of an ongoing research effort to build high-quality speech interaction datasets for next-generation conversational AI systems.
The dataset is still under active development. Data organization, validation, annotation, and quality control are continuously being improved before a stable public release.… See the full description on the dataset page: https://huggingface.co/datasets/ShiniChien/OmniEdu.OmniDistil
Overview
This dataset is being developed for training and evaluating omni-modal speech models, with a primary focus on audio-to-audio tasks. The repository is part of an ongoing research effort to build high-quality speech interaction datasets for next-generation conversational AI systems.
The dataset is still under active development. Data organization, validation, annotation, and quality control are continuously being improved before a stable public release.… See the full description on the dataset page: https://huggingface.co/datasets/ShiniChien/OmniDistil.
