datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Designed-Vocalizations-Dataset
Designed Vocalizations Dataset
Paper · Demo & audio samples
The Designed Vocalizations Dataset supports voice conversion for designed vocalizations
— monster growls, robotic voices, and other sound-designed timbres — an area left
underexplored by benchmarks that focus on natural human speech. It curates diverse raw vocal
sources (speech and animal / non-linguistic sounds) and applies professional vocal-effects
processing to produce corresponding effect-modified variants. A… See the full description on the dataset page: https://huggingface.co/datasets/NCSOFT/Designed-Vocalizations-Dataset.QWEN3-TTS-Voice-Design-100-Japanese-Female-Designed-Voices100 Japanese Female Designed Voices by Qwen3-TTS-12Hz-1.7B-VoiceDesign
Note: Contains frequent misreadings. Correct reading data is not provided.
AI Generation: The 100 styles were automatically generated by AI, so there may be some overlaps or duplicates.
voice is designed by Japanese Prompt(see styles_jp.txt)
Dataset: 300 audio clips (100 styles × 3 iterations).
Structure: design1–design3 represent each iteration. Each output is unique.
Fixes: Replaced one instance of a "complete error"… See the full description on the dataset page: https://huggingface.co/datasets/Akjava/QWEN3-TTS-Voice-Design-100-Japanese-Female-Designed-Voices.QWEN3-TTS-Voice-Design-100-Japanese-Female-Designed-Voices100 Japanese Female Designed Voices by Qwen3-TTS-12Hz-1.7B-VoiceDesign
Note: Contains frequent misreadings. Correct reading data is not provided.
AI Generation: The 100 styles were automatically generated by AI, so there may be some overlaps or duplicates.
voice is designed by Japanese Prompt(see styles_jp.txt)
Dataset: 300 audio clips (100 styles × 3 iterations).
Structure: design1–design3 represent each iteration. Each output is unique.
Fixes: Replaced one instance of a "complete error"… See the full description on the dataset page: https://huggingface.co/datasets/196vm3/QWEN3-TTS-Voice-Design-100-Japanese-Female-Designed-Voices.voice-design-bench-50voice-design-bench-50-dnsmos
Model
Mean
Min
Max
Qwen
3.262
2.480
3.619
Echo
3.217
2.158
3.588
Omni
3.190
0.988
3.629
ML-Based-Procedural-Sound-Design
ML-Based Procedural Sound Design Artifacts
This repository stores large experiment artifacts for the ML-Based-Procedural-Sound-Design project.
Contents
explosion_dataset.zip: Full explosion dataset
embeddings_630k_audioset.zip: CLAP-based embedding outputs used in the project
kmeans_results.zip: K-means clustering outputs and related result files
Notes
This repository is used for artifact storage only. Source code, evaluation figures, documentation, and the… See the full description on the dataset page: https://huggingface.co/datasets/krvieni/ML-Based-Procedural-Sound-Design.faceocultavoice_designA dataset for voice design.
~30K samples of voice design data (LAION)
