datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ovos-tts-bench-intents-for-eval-prompts
OVOS tts bench — intents-for-eval-prompts
Synthesised clips (one per prompt) predictions of the registered
OVOS Plugin Arena
tts fighters over
OpenVoiceOS/intents-for-eval.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-tts-bench-intents-for-eval-prompts.episodes
My Weird Prompts - Episode Dataset
The production record of every episode of the
My Weird Prompts podcast: the transcript, links to
the published episode, a description of the prompt that started it, and the
generation telemetry for how it was made - model, pipeline version, GPU, timings
and compute cost.
5,355 episodes. Synced daily from the production database.
from datasets import load_dataset
ds = load_dataset("My-Weird-Prompts/episodes", split="train")
Which… See the full description on the dataset page: https://huggingface.co/datasets/My-Weird-Prompts/episodes.gigaspeech-l_multi_promptslibri960h_multi_promptslibrispeech360h_multi_promptslibri500_multi_promptsadaption-music-style-prompts
This dataset is a remastered version of Reubencf/fma-labeled prepared using Adaption's Adaptive Data platform.
music_style_prompts
This dataset contains a collection of descriptive text prompts designed to generate diverse musical tracks across various genres, including pop, techno, ambient, and rock. Each entry details specific instrumentation, rhythmic patterns, atmospheric qualities, and emotional tones to guide audio synthesis. The content serves as a resource Intended for… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/adaption-music-style-prompts.moshi-on-policy-prompts-3kBased on https://huggingface.co/datasets/yhytoto12/behavior-sd
talromur3_with_prompts
Overview
talromur_3_with_prompts is a prompt-labelled corpus that can be used for fine-tuning models, such as ParlerTTS.The corpus consists of approximately 15,000 utterances, spoken by 7 named speakers in 6 different emotions (see more info here).
The dataset is an expanded version of Talromur-3: an Icelandic emotional speech corpus.We have added natural-language descriptions of utterance-level pitch, speech monotony, speech quality, reverberation, speaking rate, and emotional… See the full description on the dataset page: https://huggingface.co/datasets/atlithor/talromur3_with_prompts.patching-music-musiccaps-prompts
Activation-patching prompt pairs (MusicCaps-derived)
Paper
TADA! Tuning Audio Diffusion Models through Activation Steering — https://huggingface.co/papers/2602.11910
3246 rows of (clean, corrupted) prompt pairs derived from the MusicCaps captions by swapping feature-bearing words (e.g. violin↔trumpet, female↔male, fast↔slow) using the mapping in src/preprocess/features.py.
Features covered (21): bongos, cello, drums, fast, female, flute, happy, harmonica, jazz, male… See the full description on the dataset page: https://huggingface.co/datasets/lukasz-staniszewski/patching-music-musiccaps-prompts.videos-promptslibri_test_with_promptsoverfitting_with_promptsemilia-en-itts-prompts
Emilia EN ITTS Prompts (3 disjoint sets, with audio)
Three disjoint random subsets of 10,000 English clips each (30,000 total,
~77.5 h) sampled from amphion/Emilia-Dataset
(Emilia/EN/*.tar), for use as inference-time TTS (ITTS) generation prompts.
Each row includes the original Emilia audio (mp3, 24 kHz) plus its transcript
and metadata.
Splits
Split
Rows
set_1
10,000
set_2
10,000
set_3
10,000
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/dlion168/emilia-en-itts-prompts.
