datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
common_voice_17_0genshin-voice
Genshin Voice
Genshin Voice is a dataset of voice lines from the popular game Genshin Impact.
Hugging Face 🤗 Genshin-Voice
ModelScope Genshin-Voice
Per-speaker downloads are grouped by language and ZIP size. Browse every archive in the ZIP index.
Last update at 2026-08-13
654252 wavs
7291 without speaker (1%)
52693 without transcription (8%)
1088 without inGameFilename (0%)
Dataset Details
Dataset Description
The dataset contains voice lines… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/genshin-voice.Teachers
仓库信息
电报地址:https://t.me/vomebook ,有问题请在:https://huggingface.co/datasets/VoiceOfML/Teachers/discussions 提出。
此仓库存储导师著作:https://huggingface.co/datasets/VoiceOfML/Teachers/tree/main 。
请使用:https://voiceofml-search.hf.space/Teachers 进行文件检索(备用搜索站:https://vomebook.github.io/search/#/Teachers )。
可使用:https://voiceofml-search.hf.space/Teachers?wide=1 进行仓库内容查看(备用站:https://voiceofml-search.hf.space/Teachers?wide=1 )。
你可以仅下载指针(只有文件名的信息)
If you want to clone without large files - just their… See the full description on the dataset page: https://huggingface.co/datasets/VoiceOfML/Teachers.voicebench
License
The dataset is available under the Apache 2.0 license.
Citation
If you use the VoiceBench dataset in your research, please cite the following paper:
@article{chen2024voicebench,
title={VoiceBench: Benchmarking LLM-Based Voice Assistants},
author={Chen, Yiming and Yue, Xianghu and Zhang, Chen and Gao, Xiaoxue and Tan, Robby T. and Li, Haizhou},
journal={arXiv preprint arXiv:2410.17196},
year={2024}
}
zenless-voice
Zenless Voice
Zenless Voice is a dataset of voice lines from the popular game Zenless Zone Zero.
Hugging Face 🤗 Zenless-Voice
ModelScope Zenless-Voice
Per-speaker downloads are grouped by language and WAV count. Browse every archive in the ZIP index.
Last update at 2026-09-17, game version 3.2.0
406720 wavs
78785 without speaker (19%)
123429 without transcription (30%)
83509 without inGameFilename (21%)
Speaker archives contain 327,935 WAVs in 4,322 ZIPs. The 78,785 rows… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/zenless-voice.SovMaterials
仓库信息
电报地址:https://t.me/vomebook ,有问题请在:https://huggingface.co/datasets/VoiceOfML/SovMaterials/discussions 提出。
此仓库存储苏联资料:https://huggingface.co/datasets/VoiceOfML/SovMaterials/tree/main 。
请使用:https://voiceofml-search.hf.space/SovMaterials 进行文件检索(备用搜索站:https://vomebook.github.io/search/#/SovMaterials )。
可使用:https://voiceofml-search.hf.space/SovMaterials?wide=1 进行仓库内容查看(备用站:https://voiceofml-search.hf.space/SovMaterials?wide=1 )。
你可以仅下载指针(只有文件名的信息)
If you want to clone without… See the full description on the dataset page: https://huggingface.co/datasets/VoiceOfML/SovMaterials.portuguese-male-voice-A-datasetgenshin-voice-v3.3-mandarin
Dataset Card for Genshin Voice
Dataset Description
Dataset Summary
The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game.
Languages
The text in the dataset is in Mandarin.
Dataset Creation
Source Data
Initial Data Collection and Normalization
The data was obtained by unpacking the Genshin Impact game.
Who are the source language producers?
The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.3-mandarin.uyghur-common-voice-tts
Uyghur Common Voice TTS Dataset
A cleaned and processed Text-to-Speech (TTS) dataset for the Uyghur language, derived from Mozilla Common Voice.
Dataset Summary
Property
Value
Language
Uyghur (ug)
Total Samples
43,054
Train Samples
40,901
Validation Samples
2,153
Audio Format
WAV
Source
Mozilla Common Voice
License
CC0-1.0
Dataset Structure
/
├── train.jsonl # Training data (40,901 samples)
├── val.jsonl #… See the full description on the dataset page: https://huggingface.co/datasets/anke01/uyghur-common-voice-tts.VoiceAgentBench
VoiceAgentBench
This repository contains dataset for VoiceAgentBench, a large-scale speech benchmark introduced in “VoiceAgentBench: Are Voice Assistants Ready for Agentic Tasks?” (arXiv:2510.07978).
VoiceAgentBench is designed to evaluate end-to-end speech-based agents in realistic, tool-driven settings. Unlike prior speech benchmarks that focus on transcription, intent detection, and speech question answering, this benchmark targets agentic reasoning from speech input, requiring… See the full description on the dataset page: https://huggingface.co/datasets/krutrim-ai-labs/VoiceAgentBench.X-Voice-Dataset-Train
X-Voice Training Dataset
Overview
The X-Voice training dataset is a large-scale multilingual speech corpus curated for high-performance speech models. It provides a robust foundation for cross-lingual phonetic and prosodic modeling.
Also the train set of X-Voice Model.
Core Statistics
Total Speech Duration: 420K hours
30 languages
European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Dataset-Train.VoiceAssistant-400KGPCREducation
仓库信息
电报地址:https://t.me/vomebook ,有问题请在:https://huggingface.co/datasets/VoiceOfML/GPCREducation/discussions 提出。
此仓库存储文化大革命教材:https://huggingface.co/datasets/VoiceOfML/GPCREducation/tree/main 。
请使用:https://voiceofml-search.hf.space/GPCREducation 进行文件检索(备用搜索站:https://vomebook.github.io/search/#/GPCREducation )。
可使用:https://voiceofml-search.hf.space/GPCREducation?wide=1 进行仓库内容查看(备用站:https://voiceofml-search.hf.space/GPCREducation?wide=1 )。
你可以仅下载指针(只有文件名的信息)
If you want to clone… See the full description on the dataset page: https://huggingface.co/datasets/VoiceOfML/GPCREducation.common_voice_16_0common-voice-13-faThe Persian portion of the original CommonVoice 13 dataset at https://huggingface.co/datasets/mozilla-foundation/common_voice_13_0
Load
# Using HF Datasets
from datasets import load_dataset
dataset = load_dataset("hezarai/common-voice-13-fa", split="train")
# Using Hezar
from hezar.data import Dataset
dataset = Dataset.load("hezarai/common-voice-13-fa", split="train")
sorrel-sft-voicelaions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuningLAION's Got Talent: Generated Voice Acting Dataset
Overview
"LAION's Got Talent" is a synthetic voice acting dataset designed to offer a broad range of emotional expressions, vocal bursts, and multi-language utterances. This dataset is a component of the BUD-E project, led by LAION with support from Intel, and aims to drive forward research in context-aware and empathetic AI voice assistants.
Updated Composition
Voices and Languages
English: 11 OpenAI voices, each… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuning.starrail-voice
StarRail Voice
StarRail Voice is a dataset of voice lines from the popular game Honkai: Star Rail.
Hugging Face 🤗 StarRail-Voice
ModelScope StarRail-Voice
Last update at 2026-07-16, game version 4.4.0
403437 wavs
60164 without speaker (15%)
61375 without transcription (15%)
57869 without inGameFilename (14%)
Dataset Details
Dataset Description
The dataset contains voice lines from the game's characters in multiple languages, including Chinese… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/starrail-voice.VoiceBank-DEMAND-16kA-Historical-Learning-Data
仓库信息
如题,这是一个历史性地存在过的,而现在已经不存在的资料库的整理。
来自“revorevo.gitlab.io/mlmmlm-icu-2022/t/topic/130.html”的学习书单的资源整理。
电报地址:https://t.me/vomebook ,有问题请在:https://huggingface.co/datasets/VoiceOfML/A-Historical-Learning-Data/discussions 提出。
你可以仅下载指针(只有文件名的信息)
If you want to clone without large files - just their pointers
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/datasets/VoiceOfML/A-Historical-Learning-Data
其余仓库
仓库
链接
马列之声ebook… See the full description on the dataset page: https://huggingface.co/datasets/VoiceOfML/A-Historical-Learning-Data.common_voice_21_0_minicommon-voice-scripted-speech-26
Common Voice Scripted Speech
A row-normalized multilingual ASR dataset built from Mozilla Data Collective
Common Voice Scripted Speech. Each upstream archive is converted to appendable
parquet shards under data/<upstream_split>/, one shard per source archive and
split, with audio bytes embedded in an audio struct column.
Status
Manifest languages: 60
Languages uploaded: 18
Columns
audio (bytes, path)
sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.Omnibus
仓库信息
电报地址:https://t.me/vomebook ,有问题请在:https://huggingface.co/datasets/VoiceOfML/Omnibus/discussions 提出。
此仓库存储书纳百川专题:https://huggingface.co/datasets/VoiceOfML/Omnibus/tree/main 。
请使用:https://voiceofml-search.hf.space/Omnibus 进行文件检索(备用搜索站:https://vomebook.github.io/search/#/Omnibus )。
可使用:https://voiceofml-search.hf.space/Omnibus?wide=1 进行仓库内容查看(备用站:https://voiceofml-search.hf.space/Omnibus?wide=1 )。
你可以仅下载指针(只有文件名的信息)
If you want to clone without large files - just their… See the full description on the dataset page: https://huggingface.co/datasets/VoiceOfML/Omnibus.voicebench
License
The dataset is available under the Apache 2.0 license.
Citation
If you use the VoiceBench dataset in your research, please cite the following paper:
@article{chen2024voicebench,
title={VoiceBench: Benchmarking LLM-Based Voice Assistants},
author={Chen, Yiming and Yue, Xianghu and Zhang, Chen and Gao, Xiaoxue and Tan, Robby T. and Li, Haizhou},
journal={arXiv preprint arXiv:2410.17196},
year={2024}
}
arc-voicesamples-generatedwav2vec2_common_voice_accents_3voiceNam
Dataset Tiếng Việt (Voice Nữ)
Đây là bộ dữ liệu bao gồm file âm thanh và transcript tương ứng, được sử dụng cho việc train mô hình TTS (Text-to-Speech).
Cấu trúc dữ liệu
audio: File âm thanh (.wav)
text: Nội dung văn bản tương ứng
Cách sử dụng
from datasets import load_dataset
dataset = load_dataset("doduy1911/voiceNamNam", split="train")
# Nghe thử mẫu đầu tiên
print(dataset[0]["text"])
voice-data
Voice-Data: a curated multi-corpus voice dataset for voice–text contrastive training
voice-data is a single, globally-shuffled WebDataset that bundles several voice/speech corpora into one ready-to-train mixture for voice–text contrastive (CLAP-style) models such as VoiceCLAP. Each clip pairs 48 kHz mono FLAC audio with a natural-language text caption describing the voice — its emotion, prosody, timbre, speaking style, recording context, and speaker traits.
The distinguishing… See the full description on the dataset page: https://huggingface.co/datasets/gijs/voice-data.VoiceAssistant-400K-SLAM-Omni
VoiceAssistant-400K (Modified)
This dataset is prepared for the reproduction of SLAM-Omni.
This is a single-round English spoken dialogue training dataset. For code and usage examples, please refer to the related GitHub repository: X-LANCE/SLAM-LLM (examples/s2s)
🔧 Modifications
Data Filtering: We removed samples with excessively long data.
Speech Response Tokens: We used CosyVoice to synthesize corresponding semantic speech tokens for the speech response. These… See the full description on the dataset page: https://huggingface.co/datasets/worstchan/VoiceAssistant-400K-SLAM-Omni.laion-voice-profiles-annotated
LAION Voice Profiles — Annotated
Authors: Christoph Schuhmann and LAION.
28,212,933 utterances / 71,056 hours of synthetic English and German voice-acting speech from
500 distinct voice profiles, each driven through the same fixed matrix of 842 named acting
conditions. Every
utterance carries 40 emotion intensities, 57 perceptual voice dimensions, 4 audio-quality heads,
vocal-burst detections with timings, word-level forced alignment, MOSS audio codec tokens, a
768-d… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-annotated.
