datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
icelandic_asr
Icelandic ASR Collection
This repository collects six Icelandic speech corpora in directly loadable
Parquet form. Audio is embedded as 16 kHz mono FLAC bytes. The repository is a
convenience repackaging: the linked CLARIN-IS records and original dataset
repositories remain the canonical sources and should be cited when using the
data.
No configuration is selected by default. Choose a corpus configuration and,
for this large collection, normally choose a split explicitly.… See the full description on the dataset page: https://huggingface.co/datasets/Aalto-Speech-Synthesis/icelandic_asr.SIWIS_French_Speech_Synthesis_Database
SIWIS French Speech Synthesis Database
This README provides a concise description of the dataset, including its structure, file naming conventions, and known labeling issues. Additionally, suggestions for potential improvements are outlined in the TODO section.
The dataset is distributed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license, permitting its use for any purpose.
For more details about the database design and recording process, please refer… See the full description on the dataset page: https://huggingface.co/datasets/Aviv-anthonnyolime/SIWIS_French_Speech_Synthesis_Database.enni-child-speech-synthesislicense: mit
task_categories:
text-to-speech
automatic-speech-recognition
language:
en
tags:
speech
audio
child-speech
talkbank
size_categories:
10K<n<100K
TalkBank Child Speech Synthesis Dataset (Seed 1)
This dataset contains child speech synthesis data generated from the TalkBank FASA ENNI corpus.
Dataset Information
Number of Samples: 10032
Seed: 1
Audio Format: WAV (16kHz)
Source: TalkBank FASA ENNI
Data Structure
The dataset contains the following… See the full description on the dataset page: https://huggingface.co/datasets/jsun39/enni-child-speech-synthesis.romanian_speech_synthesis_0_8_1\
The Romanian speech synthesis (RSS) corpus was recorded in a hemianechoic chamber (anechoic walls and ceiling; floor partially anechoic) at the University of Edinburgh. We used three high quality studio microphones: a Neumann u89i (large diaphragm condenser), a Sennheiser MKH 800 (small diaphragm condenser with very wide bandwidth) and a DPA 4035 (headset-mounted condenser). Although the current release includes only speech data recorded via Sennheiser MKH 800, we may release speech data recorded via other microphones in the future. All recordings were made at 96 kHz sampling frequency and 24 bits per sample, then downsampled to 48 kHz sampling frequency. For recording, downsampling and bit rate conversion, we used ProTools HD hardware and software. We conducted 8 sessions over the course of a month, recording about 500 sentences in each session. At the start of each session, the speaker listened to a previously recorded sample, in order to attain a similar voice quality and intonation.stortinget_speech_corpus_v1.0
Dataset Card for Stortinget Speech Corpus V1.0
Overview
This is the WebDataset version of the Stortinget Speech Corpus V1.0, originally created by the National Library of Norway. We re-organize it into WebDataset format for better usability.
The Stortinget Speech Corpus (SSC) is a 5000+ hours speech dataset for weak supervision ASR created from audio andaligned proceedings text from Stortinget, the Norwegian Parliament. For more information, please refer to the original… See the full description on the dataset page: https://huggingface.co/datasets/Aalto-Speech-Synthesis/stortinget_speech_corpus_v1.0.Towards-Joint-Modeling-of-Dialogue-Response-and-Speech-Synthesis-based-on-Large-Language-Modelcommon_voice_17_0_romanian_speech_synthesiscommon_voice_16_1_romanian_speech_synthesiscommon_voice_romanian_speech_synthesisspeech-synthesisChinese_Female_Speech_Synthesis_Corpus_Live_Streaming_for_Sales_with_Multi_Styles
ID
King-TTS-241
Duration
8.56 hours
Speakers
100 People
Labeling Details
Pronunciation, Rhythm, Breath sounds marked with {hx}
Language
Chinese
Description
Two styles: Deep and uplifting; covers a variety of product categories including food, clothing, beauty, personal care, electronics, and home goods.
URL… See the full description on the dataset page: https://huggingface.co/datasets/DataoceanAI/Chinese_Female_Speech_Synthesis_Corpus_Live_Streaming_for_Sales_with_Multi_Styles.Chinese_Male_Speech_Synthesis_Corpus_Live_Streaming_for_Sales
ID
King-TTS-272
Duration
4.32 hours
Language
Chinese
URL
https://dataoceanai.com/datasets/tts/chinese-male-speech-synthesis-corpus-live-streaming-for-sales/
Chinese_Female_Speech_Synthesis_Corpus_Live_Streaming_for_Sales
ID
King-TTS-271
Duration
4.24 hours
Language
Chinese
URL
https://dataoceanai.com/datasets/tts/chinese-female-speech-synthesis-corpus-live-streaming-for-sales/
American_English_Male_Speech_Synthesis_Corpus_Gentle_and_Mature_Aged_30_40
ID
King-TTS-286
Duration
3.02 hours
Language
English
Labeled Details
Multi-emotion - Neutral, Happy, Angry, Sad, Shocked, Hateful, Scared, Shouting, Crying, Laughing, Weak
URL
https://dataoceanai.com/datasets/asr/indonesian-speech-recognition-corpus/
Nord-Parl-TTS
Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech
News
2026.01.17: 🎉 Our paper "Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech" has been accepted to ICASSP 2026! See you in Barcelona! paper
Overview
Nord-Parl-TTS is an open TTS dataset for Finnish and Swedish based on speech found in the wild.
Using recordings of Nordic parliamentary proceedings, we extract 900 hours of Finnish and 5090 hours of Swedish… See the full description on the dataset page: https://huggingface.co/datasets/Aalto-Speech-Synthesis/Nord-Parl-TTS.2-People-Korean-Natural-Conversation-Average-Tone-Speech-Synthesis-Corpus
Description
48kHz, 24bit 품질의 영어 음성 데이터셋으로, 전문 녹음 스튜디오에서 전문 성우 2명(남성 1명, 여성 1명)의 음성을 수집했습니다. 주어진 주제에 대한 즉흥 발화, 다단계 감정, 단일 감정 및 준언어적 특성(Paralinguistic Features) 등 다양한 음성 콘텐츠를 포함합니다.
텍스트, 감정 및 준언어적 특성에 대한 어노테이션을 제공하며, 음성 합성(Speech Synthesis) 등의 음성 AI 모델 개발 및 학습에 활용할 수 있습니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/tts/1540?source=hf.kr
Specifications
Format
48kHz, 24bit, 비압축 WAV, 모노 채널
Recording Environment
전문 녹음 스튜디오… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/2-People-Korean-Natural-Conversation-Average-Tone-Speech-Synthesis-Corpus.American_English_Female_Speech_Synthesis_Corpus
ID
King-TTS-287
Duration
3.51 hours
Language
English
Labeled Details
Multi-emotion - Neutral, Happy, Angry, Sad, Shocked, Hateful, Scared, Shouting, Crying, Laughing, Weak
URL
https://dataoceanai.com/datasets/tts/american-english-female-speech-synthesis-corpus-mature-aged-50-60/
common_voice_17_0_romanian_speech_synthesis
