datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikimedia-common-audio-catalanThis is a collection of Catalan-language audio with free licenses extracted from Wikimedia Commons.
License identifiers are normalized to cc-zero, cc-by-4.0,
cc-by-sa-3.0, cc-by-sa-4.0, GFDL, and PD-self.
This provides a richer alternative to Common Voice.
Characteristics of the dataset:
One or multiple speakers
Different accents
Different domain texts
761 audio files
We found this dataset useful for audio tasks such as:
Language detection
Evaluation of STT systems
New candidates are… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/wikimedia-common-audio-catalan.Chinese-LiPS
Chinese-LiPS: A Chinese audio-visual speech recognition dataset with Lip-reading and Presentation Slides
⭐ Introduction
The Chinese-LiPS dataset is a multimodal dataset designed for audio-visual speech recognition (AVSR) in Mandarin Chinese. This dataset combines speech, video, and textual transcriptions to enhance automatic speech recognition (ASR) performance, especially in educational and instructional scenarios.
🚀 Dataset Details
Total Duration:… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Chinese-LiPS.Neapolitan-Spoken-Corpus
Neapolitan Spoken Corpus (NSC)
A corpus of read Neapolitan speech for ASR evaluation, with a validated
Neapolitan–Italian lexicon, LOSO fine-tuning splits, trained LoRA adapters,
metric implementations, per-clip results, and error annotations.
This release supersedes the earlier 141-clip single-speaker version of this
repository. The earlier release corresponds to Speaker S1 of the present
corpus; the old audioData/ and transcripts.csv are replaced by
data/audio/ and… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nsc-author/Neapolitan-Spoken-Corpus.central-kurdish-tts4all
TTS4All Central Kurdish Speech Dataset
Dataset Summary
The TTS4All Central Kurdish Speech Dataset is a multi-speaker speech corpus developed for speech synthesis and speech technology research in Central Kurdish (Sorani Kurdish).
The dataset was created within the TTS4All initiative during the JSALT 2025 Workshop and provides more than 35 hours of transcribed speech from three native Central Kurdish speakers.
The corpus was designed to support:
Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-tts4all.clipquill-asr-benchmark
Measuring whisper-tiny vs whisper-base in a browser tab
Word error rate, wall-clock timing, transfer size and peak memory for two
quantised Whisper tiers running entirely client-side in a real Chrome window,
with the scripts that produced every number.
If you are building an in-browser transcription page, the two results worth
knowing before you pick a model tier:
On clean synthetic audio the two tiers tie. If that is all you test, you
will conclude the tier does not matter… See the full description on the dataset page: https://huggingface.co/datasets/sophia8888/clipquill-asr-benchmark.id-en-codeswitch-dataset-alternative
Indonesian–English Code-Switching Synthetic Speech Dataset
Synthetic speech generated for the undergraduate final project "Handling
Code-Switching in Automatic Speech Recognition for Low-Resource Language
Pairs: An Indonesian–English Case Study", School of Electrical Engineering
and Informatics, Institut Teknologi Bandung.
This dataset contains synthetic audio produced from the Indonesian–English
code-switching text corpora released in the companion repository below. It
was used… See the full description on the dataset page: https://huggingface.co/datasets/shulhaaja/id-en-codeswitch-dataset-alternative.The_Arabic_News_speech_Corpus_Dataset
Arabic News Speech Corpus Dataset
This dataset is an Arabic speech corpus that supports the development of syllable-based Arabic speech recognition using Wav2Vec-2 architecture and a 5-gram language model. It consists of Modern Standard Arabic (MSA) syllables extracted from TV news broadcasts, annotated with diacritics.
Dataset Details
Dataset Description
This corpus contains 15 hours of WAV audio recordings transcribed into diacritized Modern Standard Arabic… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimSalah/The_Arabic_News_speech_Corpus_Dataset.test4
test4
This is a merged speech dataset containing 345 audio segments from 2 source datasets.
Dataset Information
Total Segments: 345
Speakers: 7
Languages: en
Emotions: neutral, sad, angry, happy
Original Datasets: 2
Dataset Structure
Each example contains:
audio: Audio file (WAV format, 16kHz sampling rate)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged datasets)
emotion: Detected emotion (neutral, happy… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/test4.Chinese-Speech-Dataset
🎧 Chinese (Simplified) Speech Dataset
The Chinese (Simplified) speech dataset is a high-quality speech audio dataset developed to support scalable AI and machine learning solutions with diverse and structured audio data. It contains 105 hours of speech data across 700 audio files, delivered in MP3 and WAV formats, with a total size of 229 MB. This well-balanced audio dataset provides reliable voice data, featuring 54% female and 46% male speakers, with age distribution ranging from… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Chinese-Speech-Dataset.Saudilang-Code-Switch-Corpus
SCC - Saudilang Code-Switch Corpus
The National Center for Artificial Intelligence at the Saudi Data and Artificial Intelligence Authority (SDAIA), published the "SCC" dataset, which stands for "Saudilang Code-Switch Corpus”.
This dataset contains a transcription of general conversations taken from a YouTube podcast "Thmanyah" that has been transcribed by the National Center for Artificial Intelligence in SDAIA. The data features three episodes covering different domains: investment… See the full description on the dataset page: https://huggingface.co/datasets/SDAIANCAI/Saudilang-Code-Switch-Corpus.Afar-language-text-to-speech-TTS
Usage
This dataset is designed to support the development of Text-to-Speech (TTS) systems for the Afar language. It can be integrated into web applications, mobile apps, desktop software, or other platforms that require natural-sounding Afar voice synthesis or accurate spoken language recognition.
For applications involving virtual avatars or voice personas, the following culturally appropriate voice names are recommended:
Female Voices: Emeli, Hanaawi, Kareera, Laysani, Kulsuma… See the full description on the dataset page: https://huggingface.co/datasets/Charif-Ayfarah/Afar-language-text-to-speech-TTS.apple-speechanalyzer-vs-whisper-cpp-mac
Apple SpeechAnalyzer vs whisper.cpp on Mac
Four complete speech-recognition benchmark runs over the same deterministic
40-speaker LibriSpeech test-clean snapshot:
Engine
Model path
WER
CER
Repeated median post-speech latency
Repeated p95
Apple SpeechAnalyzer
progressiveTranscription on macOS 26.5
1.98%
1.02%
125–132 ms
194–201 ms
whisper.cpp server
1.8.4 · ggml-small.en
4.28%
1.79%
122–125 ms
152–161 ms
Every run completed 40/40 clips with no failures. Accuracy… See the full description on the dataset page: https://huggingface.co/datasets/researchaudio/apple-speechanalyzer-vs-whisper-cpp-mac.human-robot-conversation-korean
Human-Robot Conversation Dataset (Korean) - 660+ Hours
Dataset (Korean) contains 660+ hours of audio featuring dialogues between AI and a human in German across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of dialogues between AI… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-korean.human-robot-conversation-korean
Human-Robot Dataset
The dataset comprises 660+ hours of audio recordings across 20,000+ files for human-robot interactions in the Korean language. It captures authentic dialogues between humans and artificial conversational agents, specifically designed for training language models and advancing speech recognition systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in robotic systems and conversational AI technologies.… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-korean.crowd-recital-yi
About
This dataset was created by crowd-sourced recording sessions in Yiddish as part of the ivrit.ai Crowd Recital project.
Volunteers read on normal desktop or mobile setting Wikipedia articles while time-stamping every sentence read.
Later this data is normalized by aligning the gathered captions with the audio using Stable Whisper (See Below).
The recording project is an ongoing effort and new data will be appended to this dataset periodically as it is being generated.… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-yi.human-robot-conversation-german
Human-Robot Dataset
The dataset comprises 660+ hours of audio recordings across 20,000+ files for human-robot interactions in the German language. It captures authentic dialogues between humans and artificial conversational agents, specifically designed for training language models and advancing speech recognition systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in speech recognition, natural language processing, and… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-german.test5
test5
This is a merged speech dataset containing 1806 audio segments from 8 source datasets.
Dataset Information
Total Segments: 1806
Speakers: 47
Languages: en
Emotions: happy, neutral, sad, angry
Original Datasets: 8
Dataset Structure
Each example contains:
audio: Audio file (WAV format, 16kHz sampling rate)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged datasets)
emotion: Detected emotion (neutral… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/test5.taiwan-conversation-context-100-domains
Taiwan Conversation Context 100 Domains
Dataset Description
Taiwan Conversation Context 100 Domains 是一套以台灣日常生活情境為核心設計的雙人對話文本資料集。
本資料集包含 100 個生活領域,每個領域各有 12,000 筆對話資料,總計約 1,200,000 筆對話樣本。每筆資料皆為雙人對話格式,包含 [A][B][A][B][A][B][A][B] 共 8 個發言,也就是 4 輪來回對話。
資料以繁體中文撰寫,並針對台灣在地語境設計,適合用於:
語音生成資料前處理
Text-to-Speech, TTS
Spoken Dialogue Generation
Conversational AI
Customer Service Dialogue Modeling
Role-play Dialogue Dataset
台灣繁體中文語音模型訓練
生活情境問答模型訓練
對話式 AI 助理訓練
RAG / Agent 測試資料… See the full description on the dataset page: https://huggingface.co/datasets/Ethan615/taiwan-conversation-context-100-domains.human-robot-conversation-english
Human-Robot Dataset
The dataset comprises 660+ hours of English speech across 20,000+ audio files featuring human-robot interactions between AI and humans. It is designed for research in conversational agents, focusing on various speech recognition methods, primarily aimed at advancing language models and machine learning applications.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in robotic systems and conversational AI… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-english.commonvoicebadini
Northern Kurdish (Arabic Script) ASR Dataset
Dataset Description
Northern Kurdish is the most widely spoken variant of the Kurdish language and is used across all parts of Kurdistan. Although it is mainly written today in the Latin script, it was historically written in the Arabic script. The Arabic script is still used for this dialect in Southern Kurdistan, particularly in the Duhok province of the Kurdistan Regional Government (KRG).Similarly, the primary writing… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/commonvoicebadini.CommonVoices20_ro
Common Voices Corpus 20.0 (Romanian)
Common Voices is an open-source dataset of speech recordings created by
Mozilla to improve speech recognition technologies.
It consists of crowdsourced voice samples in multiple languages, contributed by volunteers worldwide.
Challenges: The raw dataset included numerous recordings with incorrect transcriptions
or those requiring adjustments, such as sampling rate modifications, conversion to .wav format, and other refinements
essential… See the full description on the dataset page: https://huggingface.co/datasets/TransferRapid/CommonVoices20_ro.human-robot-conversation-english
Human-Robot Conversation Dataset (English) - 660+ Hours
Dataset (English) contains 660+ hours of audio featuring dialogues between AI and a human in English across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of dialogues between… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-english.Cebuano-Speech-Dataset
🎧 Cebuano Speech Dataset
The Cebuano Speech Dataset is a high-quality speech audio dataset designed to deliver structured and diverse audio data for AI-powered voice applications. It includes 108 hours of audio data distributed across 807 files, provided in MP3 and WAV formats, with a total size of 135 MB. This well-organized audio dataset ensures balanced voice data, with 49% female and 51% male speakers, and a broad age range from 18 to 50+ years. The dataset language is Cebuano… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Cebuano-Speech-Dataset.test6
test6
This is a merged speech dataset containing 1994 audio segments from 2 source datasets.
Dataset Information
Total Segments: 1994
Speakers: 3
Languages: en
Emotions: neutral, negative_surprise, positive_surprise, distress, relief, contentment, adoration, interest, confusion, happy, sadness, triumph, fear, disappointment, awe, realization, angry
Original Datasets: 2
Dataset Structure
Each example contains:
audio: Audio file (WAV format, 16kHz sampling… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/test6.crowd-whatsapp-yi
About
This dataset was created by crowd-sourced Whatsapp voice recordings in Yiddish as part of the ivrit.ai project.
Volunteers read a message sent to them from a predefined set of messages, recording themselves using Whasapp voice message sent to the collecting bot.
Later this data is normalized by aligning the captions with the audio using Stable Whisper (See Below).
The recording project is an ongoing effort and new data will be appended to this dataset periodically as it is… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-whatsapp-yi.last
last
This is a merged speech dataset containing 345 audio segments from 2 source datasets.
Dataset Information
Total Segments: 345
Speakers: 7
Languages: en
Emotions: neutral, angry, happy, sad
Original Datasets: 2
Dataset Structure
Each example contains:
audio: Audio file (WAV format, 16kHz sampling rate)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged datasets)
emotion: Detected emotion… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/last.Chichewa-Speech-Dataset
🎧 Chichewa Speech Dataset
The Chichewa Speech Dataset is a high-quality speech audio dataset designed to provide structured and scalable audio data for AI-driven voice technologies. It includes 94 hours of audio data across 740 files, delivered in MP3 and WAV formats, with a total size of 272 MB. This well-organized audio dataset ensures balanced and diverse voice data, with 52% female and 48% male speakers, and a wide age distribution from 18 to 50+ years. The dataset language is… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Chichewa-Speech-Dataset.tr123
tr123
This is a merged speech dataset containing 11930 audio segments from 24 source datasets.
Dataset Information
Total Segments: 11930
Speakers: 69
Languages: tr
Emotions: sad, neutral, happy, angry
Original Datasets: 24
Dataset Structure
Each example contains:
audio: Audio file (WAV format, 16kHz sampling rate)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged datasets)
emotion: Detected… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/tr123.testtr12
testtr12
This is a merged speech dataset containing 165 audio segments from 2 source datasets.
Dataset Information
Total Segments: 165
Speakers: 8
Languages: en
Emotions: happy, angry, neutral
Original Datasets: 2
Dataset Structure
Each example contains:
audio: Audio file (WAV format, 16kHz sampling rate)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged datasets)
emotion: Detected emotion… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/testtr12.human-robot-conversation-german
Human-Robot Conversation Dataset (German) - 660+ Hours
Dataset (German) contains 660+ hours of audio featuring dialogues between AI and a human in German across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of dialogues between AI… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-german.
