datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
common-voice-scripted-speech-26
Common Voice Scripted Speech
A row-normalized multilingual ASR dataset built from Mozilla Data Collective
Common Voice Scripted Speech. Each upstream archive is converted to appendable
parquet shards under data/<upstream_split>/, one shard per source archive and
split, with audio bytes embedded in an audio struct column.
Status
Manifest languages: 60
Languages uploaded: 18
Columns
audio (bytes, path)
sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.voice-code-bench
VoiceCodeBench
VoiceCodeBench is a test-only benchmark for evaluating whether automatic
speech recognition (ASR) systems preserve exact structured values in English
workplace speech.
Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition
The benchmark targets cases where a transcript is software input: callback
numbers, email addresses, command-line flags, file paths, URLs, account
identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/voice-code-bench.arknights_voices_zh
ZH Voice-Text Dataset for Arknights Waifus
This is the ZH voice-text dataset for arknights playable characters. Very useful for fine-tuning or evaluating ASR/ASV models.
Only the voices with strictly one voice actor is maintained here to reduce the noise of this dataset.
12431 records, 25.9 hours in total. Average duration is approximately 7.49s.
id
char_id
voice_actor_name
voice_title
voice_text
time
sample_rate
file_size
filename
mimetype
file_url
char_106_franka_CN_001… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/arknights_voices_zh.common_voice_26_0_de
Mozilla Common Voice 26.0 - German (IPA & Clean Validated Subset)
Repacking version of Common Voice 26.0 German officialy published by Mozilla Data Collective, following Hugging Face Parquet Shards standard, with feature for listening to audio directly on the Web Hub, and the addition of a data column for the IPA transcription of each sentence.
📊 Dataset parameters
Origin: Mozilla Common Voice 26.0 (version 18/06/2026).
Data amount (Validated): 950,877 MP3 audio… See the full description on the dataset page: https://huggingface.co/datasets/q1805/common_voice_26_0_de.VoiceCommandAudioThis is mainly used for fine tune "VoiceCommand" a speech congnition MOD dedicated for SilentHunter game series
APAC-Egocentric-Residential-Voiceover
APAC Egocentric Residential (with Voiceover)
Ten narrated first-person recordings of household chores, each shipping the original capture with spoken voiceover, a burned-in caption render, WebVTT captions, and an ASS annotation track.
This is the only release in the HumynLabs egocentric collection that carries audio narration — the wearer describes each action as they perform it, and the captions align that speech to the video.
Preview: 45 s from the cooking sample, captioned… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/APAC-Egocentric-Residential-Voiceover.azurlane_voices_jp
JP Voice-Text Dataset for Azur Lane Waifus
This is the JP voice-text dataset for azur lane playable characters. Very useful for fine-tuning or evaluating ASR/ASV models.
Only the voices with strictly one voice actor is maintained here to reduce the noise of this dataset.
30160 records, 75.8 hours in total. Average duration is approximately 9.05s.
id
char_id
voice_actor_name
voice_title
voice_text
time
sample_rate
file_size
filename
mimetype
file_url… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/azurlane_voices_jp.arknights_voices_jp
JP Voice-Text Dataset for Arknights Waifus
This is the JP voice-text dataset for arknights playable characters. Very useful for fine-tuning or evaluating ASR/ASV models.
Only the voices with strictly one voice actor is maintained here to reduce the noise of this dataset.
10905 records, 26.3 hours in total. Average duration is approximately 8.7s.
id
char_id
voice_actor_name
voice_title
voice_text
time
sample_rate
file_size
filename
mimetype
file_url
char_427_vigil_CN_001… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/arknights_voices_jp.fgo_voices_jp
JP Voice-Text Dataset for FGO Waifus
This is the JP voice-text dataset for FGO playable characters. Very useful for fine-tuning or evaluating ASR/ASV models.
Only the voices with strictly one voice actor is maintained here to reduce the noise of this dataset.
30800 records, 66.4 hours in total. Average duration is approximately 7.76s.
id
char_id
voice_actor_name
voice_title
voice_text
time
sample_rate
file_size
filename
mimetype
file_url
char_1_SV1_0_对话12
1
高桥李依
对话 12… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/fgo_voices_jp.voice-code-bench
VoiceCodeBench
VoiceCodeBench is a test-only benchmark for evaluating whether automatic
speech recognition (ASR) systems preserve exact structured values in English
workplace speech.
Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition
The benchmark targets cases where a transcript is software input: callback
numbers, email addresses, command-line flags, file paths, URLs, account
identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/yunqi1766/voice-code-bench.arknights_voices_en
EN Voice-Text Dataset for Arknights Waifus
This is the EN voice-text dataset for arknights playable characters. Very useful for fine-tuning or evaluating ASR/ASV models.
Only the voices with strictly one voice actor is maintained here to reduce the noise of this dataset.
8216 records, 16.1 hours in total. Average duration is approximately 7.05s.
id
char_id
voice_actor_name
voice_title
voice_text
time
sample_rate
file_size
filename
mimetype
file_url
char_214_kafka_CN_001… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/arknights_voices_en.girlsfrontline_voices_jp
JP Voice-Text Dataset for Girls Front Line Waifus
This is the JP voice-text dataset for girls front line playable characters. Very useful for fine-tuning or evaluating ASR/ASV models.
Only the voices with strictly one voice actor is maintained here to reduce the noise of this dataset.
12508 records, 20.9 hours in total. Average duration is approximately 6.01s.
id
char_id
voice_actor_name
voice_title
voice_text
time
sample_rate
file_size
filename
mimetype
file_url… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/girlsfrontline_voices_jp.Trump_Voice_Dataset
Trump Voice Dataset
This dataset contains audio clips of Donald Trump's speech from the World Economic Forum (WEF) 2018, paired with their corresponding transcriptions. The dataset is designed for text-to-speech (TTS) and speech recognition tasks.
Dataset Description
Dataset Summary
The Trump Voice Dataset consists of 20 audio samples (10 train, 10 test) extracted from Donald Trump's speech at the World Economic Forum 2018. Each audio clip is approximately 10… See the full description on the dataset page: https://huggingface.co/datasets/Sakchham19/Trump_Voice_Dataset.arknights_voices_kr
KR Voice-Text Dataset for Arknights Waifus
This is the KR voice-text dataset for arknights playable characters. Very useful for fine-tuning or evaluating ASR/ASV models.
Only the voices with strictly one voice actor is maintained here to reduce the noise of this dataset.
9996 records, 23.1 hours in total. Average duration is approximately 8.34s.
id
char_id
voice_actor_name
voice_title
voice_text
time
sample_rate
file_size
filename
mimetype
file_url
char_4046_ebnhlz_CN_001… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/arknights_voices_kr.tts-conversational-voice-20000h
TTS Voice Dataset
20,000 hours of high-fidelity 48kHz conversational audio across 30+ global, regional, and underrepresented languages, built for text-to-speech, voice cloning, and multilingual speech AI.
This repository contains the full technical specification, annotation schema, and sample metadata files (Parquet). The production dataset is rights-cleared and delivered directly to buyers. Request access to see the full schema and get real audio samples.
Overview… See the full description on the dataset page: https://huggingface.co/datasets/Datoric/tts-conversational-voice-20000h.zamai-pashto-voice2voice
ZamAI Pashto Voice2Voice
This dataset contains Pashto voice-to-voice preparation metadata for speech and translation experiments. It focuses on Pashto speech records, dialect information, transcript text, and a small viewer-ready sample manifest.
Configs
from datasets import load_dataset
metadata = load_dataset("ZamAI-Pashto/zamai-pashto-voice2voice", "metadata")
sample = load_dataset("ZamAI-Pashto/zamai-pashto-voice2voice", "viewer_sample")
Files… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-voice2voice.audio-speech-realtime-voice-agents-2026
🎙️ Audio, Speech Foundation Models & Real-Time Voice Agents Dataset (2026 Edition)
A structured research dataset featuring 1,722 domain-verified research papers and 298 official code repositories focused on Full-Duplex Speech-to-Speech LLMs, Real-Time Voice Agents (<200ms Latency), Zero-Shot TTS, Voice Cloning, OpenAI Whisper-v3, Neural Audio Codecs (EnCodec/DAC/SNAC), and Generative Music (2023–2026).
Built with Universal Scientific Engine V18.1 Diamond, providing 48 schema… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/audio-speech-realtime-voice-agents-2026.zamai-pashto-voice2voice
ZamAI Pashto Voice2Voice
Languages: psLicense: cc-by-4.0Task categories: automatic-speech-recognition, audio-to-audioSize categories: n<1K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for automatic-speech-recognition, audio-to-audio tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/zamai-pashto-voice2voice")
print(dataset)
Configs
default: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/zamai-pashto-voice2voice.voicevox-voice-corpus-metadata
VOICEVOX synthetic Japanese speech corpus
This metadata derivative documents ayousanz/voicevox-voice-corpus at
39dff6b254bf3118accad04446238ae3147156a6.
Audio payloads were not downloaded during this metadata audit.
Verified repository inventory
Corpus
Voice/style directories
WAV files
WAV bytes
ROHAN-corpus
87
400,200
90.16 GB
ita-corpus
87
36,888
6.64 GB
tsukuyomi-chan-corpus
87
8,700
3.07 GB
Total WAV paths: 445,788
Total WAV bytes: 99.87… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/voicevox-voice-corpus-metadata.default_voices_chunked_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/default_voices_chunked
Aligned dataset: instinct-org/default_voices_chunked_nfa_aligned
Rows: 134236 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked_nfa_aligned.
