datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RVCBench
RVCBench
RVCBench is a benchmark dataset for studying robustness in voice cloning, text-to-speech, speaker privacy, audio protection, adversarial audio perturbations, and related audio generation pipelines.
Dataset page: https://huggingface.co/datasets/Nanboy/RVCBench
Code repository: https://github.com/Nanboy-Ronan/RVCBench
Paper: https://arxiv.org/abs/2602.00443
RVCBench is designed for evaluating how modern voice cloning (VC), TTS, and audio generation systems behave under… See the full description on the dataset page: https://huggingface.co/datasets/Nanboy/RVCBench.navigation-corpus-dagbani-speech
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Dag Speech Segments (sentence splitting)
52799 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/navigation-corpus-speech-full-dagbani
Full-file CTC forced… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/navigation-corpus-dagbani-speech.navigation-corpus-twi-speech
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi Speech Segments (sentence splitting)
52562 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/navigation-corpus-speech-full-twi
Full-file CTC forced alignment… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/navigation-corpus-twi-speech.African_voices_naija
🇳🇬 WaZoBiaSpeech: 1,000+ Hour Nigerian Pidgin (pcm) Corpus
Version: 30 Nov 2025
NOTE: This dataset is subject to regular Updates, corrections, and expansions. Please check this repository regularly for the latest release.
🌍 Dataset Overview
WaZoBiaSpeech is a large-scale, high-quality, fully transcribed speech dataset for Nigerian Pidgin (pcm). This corpus is designed to accelerate the development of speech technology in African contexts, promoting… See the full description on the dataset page: https://huggingface.co/datasets/Africanvoice/African_voices_naija.ne-asr-dataset-nag-aug
NE ASR Augmented Dataset -- Nagamese (nag)
Augmented automatic speech recognition dataset for Nagamese (nag),
a Assamese-based creole language spoken in Nagaland, India.
Source
Augmented from sulabhkatiyar/ne-asr-nag
(original transcribed speech data from the ARTPARK-IISc Vaani project).
Language Information
Property
Value
Language
Nagamese
ISO 639-3
nag
Family
Assamese-based creole
Region
Nagaland, India
Tonal
No
Tier
D (23.76h… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-asr-dataset-nag-aug.nasle-mana-clean-chunked-30s-avasanj
Nasl-e-Mana Clean Persian Speech — corrected 30-second chunks
Corrected, provenance-preserving audio chunks collected from the Nasl-e-Mana magazine website, generated on 2026-08-30. This release supersedes the earlier unreliable proportional-mapping chunk export; that older release was not used here.
Splits
Split
Rows
Audio
Columns
labeled
4,981
41.41 hours
audio, label
to_transcribe
11,127
92.72 hours
audio
The labeled split contains the… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean-chunked-30s-avasanj.Japanese-Eroge-Voice-V2
Japanese-Eroge-Voice-V2
Description
This is the successor to the Japanese-Eroge-Voice dataset. It consists of a significantly larger collection of audio-transcription pairs extracted from Japanese eroge (adult games).
Note on Versioning: There is no overlap between this dataset (V2) and the previous version. All audio clips and transcriptions in V2 are distinct from those in the original version, providing entirely new data for research.
This version (V2) expands the… See the full description on the dataset page: https://huggingface.co/datasets/NandemoGHS/Japanese-Eroge-Voice-V2.global-news-radio-30s
Global News Radio Dataset
Multilingual news radio recordings from 51 languages across 42 countries.
Recordings
51
Total audio
1500 min (25.0 h)
Format
MP3 16kHz mono 64kbps
Parquet shards
11
Languages
51
Countries
42
Size
687 MB
Languages
Amharic, Arabic, Bashkir, Basque, Belarusian, Bengali, Brazilian Portuguese,Portugues Do Brasil,Português Brasil, Catalan, Croatian, Czech, Danish, Dutch, English, Estonian, Faroese, Finnish, Flemish… See the full description on the dataset page: https://huggingface.co/datasets/NathanRoll/global-news-radio-30s.ghana-named-entities-tts-twi
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Ghana Named Entities TTS — Twi
A Twi-language speech dataset built from descriptions of Ghana named entities
(people, places, organisations, and concepts). Each audio clip is a synthesised
reading of a passage that describes several… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-named-entities-tts-twi.naturalvoice_737h_48k
NaturalVoices — Sidon-Restored, UTMOS-Filtered (48 kHz)
High-quality English speech derived from NaturalVoices_VC_870h (JHU SmileLab), restored
with Sidon v0.1 and kept only
where restoration measurably improved perceptual quality (UTMOS gate). Each clip ships
with rich per-utterance metadata (transcript, speaker age/gender, speaking rate, emotion,
and quality scores) so it is ready for TTS / voice-cloning / ASR / paralinguistic research.
494,903 clips · 736.7 hours (clips ≥… See the full description on the dataset page: https://huggingface.co/datasets/PleasedPenguin/naturalvoice_737h_48k.nasle-mana-clean-chunked-30s
Nasl-e-Mana Clean Speech Corpus — Sentence-Safe 30s Chunks
Training-oriented WAV chunks derived from the public Nasl-e-Mana magazine audio corpus. Chunks target approximately 30 seconds and are cut at detected acoustic pauses; the labeled configuration additionally assigns only complete source-text sentences to each chunk.
Configuration
Rows
Columns
Meaning
labeled (train/)
9,886
audio, label
Sentence-grouped text/audio pairs from duration-compatible source-text… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean-chunked-30s.Japanese-Eroge-Voice
Japanese-Eroge-Voice
Description
This dataset contains pairs of audio data and corresponding transcriptions extracted from Japanese eroge (adult games) that I have personally purchased. The transcriptions are generated using the litagin/anime-whisper model.
Preprocessing Steps
The raw audio data has undergone the following preprocessing steps:
Loudness Normalization:
Audio loudness is normalized using ffmpeg's 2-pass loudnorm filter to target parameters of… See the full description on the dataset page: https://huggingface.co/datasets/NandemoGHS/Japanese-Eroge-Voice.navigation-corpus-ewe-speech
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Ewe Speech Segments (sentence splitting)
49348 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/navigation-corpus-speech-full-ewe
Full-file CTC forced alignment… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/navigation-corpus-ewe-speech.SloPalSpeech
SloPalSpeech
This dataset contains aligned and segmented Slovak speech–text pairs sourced from official plenary session recordings of the Slovak National Council (Národná rada Slovenskej republiky).It was prepared by collecting raw audio from MediaPortál NR SR and matching it with official full-text transcripts from the Joint Czech and Slovak Digital Parliamentary Library.A custom alignment and filtering pipeline segmented the recordings into short clips (≤30 seconds) with their… See the full description on the dataset page: https://huggingface.co/datasets/NaiveNeuron/SloPalSpeech.reprocessed_singapore_national_speech_corpus
Dataset Card for Reprocessed National Speech Corpus
NOTE: This is an Reprocessed version KaraKaraWitch from Recursal.The official download can be found here.
Dataset Details
Dataset Description
Dataset Description:
The National Speech Corpus (NSC) is the first large-scale Singapore English corpus, sponsored by the Info-communications and Media Development Authority (IMDA) of Singapore. The objective is to serve as a primary resource of open speech data for… See the full description on the dataset page: https://huggingface.co/datasets/recursal/reprocessed_singapore_national_speech_corpus.navigation-corpus-speech-full-dagbani
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Ghana TTS Navigation Corpus — Dagbani
Synthetic speech dataset for navigation.
Structure
audio/ – all .wav audio files
text/ – matching .txt files with transcriptions
metadata.csv – full metadata table
ne-asr-dataset-nag
Nagamese (nag) — ASR dataset
A small Nagamese (nag) speech-to-text dataset for automatic speech recognition
(ASR) of a low-resource North-East India language. Each example pairs a short audio
clip with its Romanized (Latin-script) transcript.
Source
Derived from the ARTPARK-IISc Vaani project (https://vaani.iisc.ac.in/)
Splits
Split
Samples
train
12,862
validation
1,532
test
1,717
Data fields
Each example has:… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-asr-dataset-nag.navigation-corpus-dagbani-speech
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Dag Speech Segments (sentence splitting)
52799 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/navigation-corpus-speech-full-dagbani
Full-file CTC forced… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/navigation-corpus-dagbani-speech.navigation-corpus-twi-speech
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi Speech Segments (sentence splitting)
52562 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/navigation-corpus-speech-full-twi
Full-file CTC forced alignment… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/navigation-corpus-twi-speech.vietspeech500G of vietnamese speech corpus from Youtube
transcription-corpus
UN Transcription Corpus
Two splits of UN meeting audio paired with official verbatim records.
Splits
sessions — Whole meeting sessions (SC + GA plenary)
One row per meeting. Audio from UN Web TV, verbatim records from documents.un.org.
Column
Description
symbol
UN document symbol, e.g. S/PV.9826
webtv_url
URL on UN Web TV
duration_ms
Session duration in milliseconds
num_speakers
Number of speaker turns in the verbatim record
audio_floor
Floor… See the full description on the dataset page: https://huggingface.co/datasets/united-nations/transcription-corpus.navigation-corpus-speech-full-twi
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Ghana TTS Navigation Corpus — Twi
Synthetic speech dataset for navigation.
Structure
audio/ – all .wav audio files
text/ – matching .txt files with transcriptions
metadata.csv – full metadata table
naturalvoice_737h_16k
NaturalVoices — Sidon-Restored, UTMOS-Filtered (16 kHz)
High-quality English speech derived from NaturalVoices_VC_870h (JHU SmileLab), restored
with Sidon v0.1 and kept only
where restoration measurably improved perceptual quality (UTMOS gate). Each clip ships
with rich per-utterance metadata (transcript, speaker age/gender, speaking rate, emotion,
and quality scores) so it is ready for TTS / voice-cloning / ASR / paralinguistic research.
494,903 clips · 736.7 hours (clips ≥… See the full description on the dataset page: https://huggingface.co/datasets/PleasedPenguin/naturalvoice_737h_16k.nasle-mana-clean
Nasl-e-Mana Speech Corpus
Clean, playable Persian speech audio collected from the public Nasl-e-Mana magazine WordPress site. The export contains two intentionally different collections:
Split
Rows
Columns
Meaning
labeled configuration (train/)
809
audio, label
Audio with recovered article text. These clips were identified as the consistent female source-text voice and are kept together.
to_transcribe configuration (to_transcribe/)
626
audio
Playable audio for which… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean.africanvoices-naija-batch1-summary
African Voices Naija Train Metadata Summary
This dataset contains a compact summary of metadata for the Naija training split, provided as CSV tables for inspection and analysis.
Files included:
batch_summary.csv
domain_distribution.csv
The repository contains metadata summaries only and does not include raw audio.
voice-Viet-Nam
🇻🇳 Voice Dataset Viet Nam
📝 Giới thiệu (Dataset Description)
Đây là kho dữ liệu âm thanh tiếng Việt mã nguồn mở (Voice Dataset Viet Nam). Bộ dữ liệu này được thu thập và cấu trúc theo chuẩn AudioFolder của Hugging Face, phục vụ cho các bài toán:
Automatic Speech Recognition (ASR): Nhận dạng giọng nói.
Text-to-Speech (TTS): Tổng hợp tiếng nói.
Voice Cloning: Huấn luyện mô hình nhân bản giọng nói.
📊 Thống kê dữ liệu (Statistics)
Thuộc tính… See the full description on the dataset page: https://huggingface.co/datasets/dinhlam2210/voice-Viet-Nam.English_Natural_Conversation_ASR_STT
BoxlyX English Natural Conversation Sample Dataset (ASR/STT)
📌 Overview
This repository contains high-fidelity, studio-recorded English natural conversation samples designed for training and benchmarking advanced Automatic Speech Recognition (ASR) and Speech-to-Text (STT) models.
This dataset is a curated public sample provided by BoxlyX AI Solution, showcasing our end-to-end capabilities in premium audio data generation, multi-speaker recording environment… See the full description on the dataset page: https://huggingface.co/datasets/BoxlyX/English_Natural_Conversation_ASR_STT.global-news-radio-debug
Global News Radio Dataset (1 hour per station)
Every news radio station from the Radio Browser API, recorded for 1 hour each.
Attempted
3037
Successful
2553
Failed
484
Total audio
21 hours
Parquet shards
256
Size
0.6 GB
Format
MP3 16kHz mono 64kbps
Usage
from datasets import load_dataset
ds = load_dataset("NathanRoll/global-news-radio-debug", streaming=True)
for sample in ds["train"]:
print(sample["station_name"], sample["language"]… See the full description on the dataset page: https://huggingface.co/datasets/NathanRoll/global-news-radio-debug.navigation-corpus-speech-full-ewe
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Ghana TTS Navigation Corpus — Ewe
Synthetic speech dataset for navigation.
Structure
audio/ – all .wav audio files
text/ – matching .txt files with transcriptions
metadata.csv – full metadata table
narrated-audiobooks-brsbs
Narrated Bashkir Audiobooks — BRSBS (Bashkorttele)
≈ 78.7 hours of human-narrated audiobooks, primarily in the Bashkir language, drawn from public-domain literary works and folk epics. Recorded as accessible "talking books" by the Bashkir Republican Special Library for the Blind (BRSBS) and released for language preservation and AI/ML research.
🌐 Languages of this card: English · Башҡортса · Русский
This dataset is part of a larger series published under the Bashkorttele… See the full description on the dataset page: https://huggingface.co/datasets/bashkorttele/narrated-audiobooks-brsbs.
