datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
genshin-voice
Genshin Voice
Genshin Voice is a dataset of voice lines from the popular game Genshin Impact.
Hugging Face 🤗 Genshin-Voice
ModelScope Genshin-Voice
Per-speaker downloads are grouped by language and ZIP size. Browse every archive in the ZIP index.
Last update at 2026-08-13
654252 wavs
7291 without speaker (1%)
52693 without transcription (8%)
1088 without inGameFilename (0%)
Dataset Details
Dataset Description
The dataset contains voice lines… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/genshin-voice.genshin-voice-v3.3-mandarin
Dataset Card for Genshin Voice
Dataset Description
Dataset Summary
The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game.
Languages
The text in the dataset is in Mandarin.
Dataset Creation
Source Data
Initial Data Collection and Normalization
The data was obtained by unpacking the Genshin Impact game.
Who are the source language producers?
The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.3-mandarin.arc-voicesamples-generatedmusic_genres
Dataset Card for "music_genres"
More Information needed
genshin-voice-v3.5-mandarin
Dataset Card for Genshin Voice
Dataset Description
Dataset Summary
The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game.
Languages
The text in the dataset is in Mandarin.
Dataset Creation
Source Data
Initial Data Collection and Normalization
The data was obtained by unpacking the Genshin Impact game.
Who are the source language producers?
The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.5-mandarin.gtzan-genrecommon-voice-17-en-age-gender-accentEmilia-dataset-french-with-gender
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/AdrienB134/Emilia-dataset-french-with-gender.common-voice-17-en-age-genderfma-genre-classification
FMA Genre Classification Dataset
The FMA Genre Classification Dataset is a subset of the Free Music Archive (FMA), containing audio samples and genre labels for music classification tasks. This version uses the "small" subset of FMA, which contains 8,000 tracks of 30 seconds each, evenly distributed across 8 genres.
Dataset Description
Dataset Summary
This dataset consists of 8,000 audio tracks from the Free Music Archive (FMA), each 30 seconds in length… See the full description on the dataset page: https://huggingface.co/datasets/rpmon/fma-genre-classification.commonvoice_train_gender_accent_16k
Dataset Card for "commonvoice_train_gender_accent_16k"
More Information needed
genshin-impact-voices
Genshin Impact — Voice Lines (Multi-Language)
An archive of character voice data extracted from Genshin Impact (原神), repackaged as Parquet shards per audio language.
Dataset Summary
Field
Value
Game
Genshin Impact (原神)
Publisher
HoYoverse / miHoYo Co., Ltd.
Languages
中文 (zh), 日本語 (ja), English (en), 한국어 (ko)
Game version
6.3
Source format
WAV + sidecar transcripts (.lab / .txt)
Distribution format
Apache Parquet (zstd), ~500 MiB audio per shard… See the full description on the dataset page: https://huggingface.co/datasets/ultemica/genshin-impact-voices.Genshin-Voice-JaMoCha-Generation-on-MoChaBench-VisualizerThis is just a Visualizer. Refer to this GitHub repo for detailed usage instructions: 🔗MoChaBench.
MoChaBench
MoCha is a pioneering model for Dialogue-driven Movie Shot Generation.
| 🌐Project Page | 📖Paper | 🔗Github | 🤗Demo|
We introduce our evaluation benchmark "MoChaBench", as described in Section 4.3 of the MoCha Paper.
MoChaBench is tailored for Dialogue-driven Movie Shot Generation — generating movie shots from a combination of speech and text(speech + text → video).
It… See the full description on the dataset page: https://huggingface.co/datasets/CongWei1230/MoCha-Generation-on-MoChaBench-Visualizer.genshin-voice-v3.4-mandarin
Dataset Card for Genshin Voice
Dataset Description
Dataset Summary
The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game.
Languages
The text in the dataset is in Mandarin.
Dataset Creation
Source Data
Initial Data Collection and Normalization
The data was obtained by unpacking the Genshin Impact game.
Who are the source language producers?
The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.4-mandarin.Hy-Generated-audio-data-with-cv20.0
Hy-Generated Audio Data with CV20.0
This dataset provides Armenian speech data consisting of both real and generated audio clips.
The train, test, and eval splits are derived from the Common Voice 20.0 Armenian dataset.
The generated split contains 100,000 high-quality clips synthesized using a fine-tuned F5-TTS model, covering 404 equal distribution of synthetic voices.
📊 Dataset Statistics
Split
# Clips
Duration (hours)
train
9,300
13.53
test
5,818… See the full description on the dataset page: https://huggingface.co/datasets/ErikMkrtchyan/Hy-Generated-audio-data-with-cv20.0.Hy-Generated-audio-data-2
Hy-Generated Audio Data 2
This dataset provides Armenian speech data consisting of generated audio clips and is addition to this dataset.
The generated split contains 137,419 high-quality clips synthesized using a fine-tuned F5-TTS model, covering 404 equal distribution of synthetic voices.
📊 Dataset Statistics
Split
# Clips
Duration (hours)
generated
137,419
173.76
Total duration: ~173 hours
🛠️ Loading the Dataset
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/ErikMkrtchyan/Hy-Generated-audio-data-2.ATC_TTS_generatedATC_generated_realGenshin-Voice-JaKimi-Audio-GenTest
Kimi-Audio-Generation-Testset
Dataset Description
Summary: This dataset is designed to benchmark and evaluate the conversational capabilities of audio-based dialogue models. It consists of a collection of audio files containing various instructions and conversational prompts. The primary goal is to assess a model's ability to generate not just relevant, but also appropriately styled audio responses.
Specifically, the dataset targets the model's proficiency in:… See the full description on the dataset page: https://huggingface.co/datasets/moonshotai/Kimi-Audio-GenTest.mmau_generation_test
MMAU-Mini Text-to-Audio SFT (realistic, v1)
Supervised fine-tuning data for open-ended text-to-audio generation. Each example
pairs a natural-language user request with an assistant turn that reasons about the
request and then produces the target audio. Built by reverse-engineering plausible user
requests from the rich captions of the MMAU-Mini test set.
995 examples across three categories: music (331), sound (331), speech (333).
Realistic prompts only (requests that directly… See the full description on the dataset page: https://huggingface.co/datasets/JinchuanTian/mmau_generation_test.genshin-voice-zhgenshin_ch_10npc
Dataset Card for "genshin_ch_10npc"
More Information needed
gtzan-music-genre-dataset
GTZAN Music Genre Dataset
The GTZAN Music Genre Dataset is a collection of 1000 audio tracks each 30 seconds long. It contains 10 genres, each represented by 100 tracks. The tracks are all 22050Hz Mono 16-bit audio files in .wav format.
Overview
This dataset was created in 2002 by George Tzanetakis and Perry Cook for research in automatic music genre classification. It has become a standard benchmark dataset in the music information retrieval (MIR) community.… See the full description on the dataset page: https://huggingface.co/datasets/storylinez/gtzan-music-genre-dataset.music_genres_small
Dataset Card for "music_genres_small"
More Information needed
MARBLEGenreClassification_MTG-Genre-Fold1
Dataset Card for "MARBLEGenreClassification_MTG-Genre-Fold1"
More Information needed
genshin_voice_longGenshin_Impact_RaidenShogun_Voice_koreangenshin-voice
Genshin Voice
Genshin Voice is a dataset of voice lines from the popular game Genshin Impact.
Hugging Face 🤗 Genshin-Voice
Last update at 2025-04-22
424011 wavs
40907 without speaker (10%)
40000 without transcription (9%)
10313 without inGameFilename (2%)
Dataset Details
Dataset Description
The dataset contains voice lines from the game's characters in multiple languages, including Chinese, English, Japanese, and Korean.
The voice lines are spoken… See the full description on the dataset page: https://huggingface.co/datasets/DeQuackDealer/genshin-voice.
