dramabox
Datasets
All datasets matching “dramabox”ears-dramabox
EARS — DramaBox Training Data
Multi-speaker DramaBox training dataset with 16,377 samples from the EARS (Expressive Anechoic Recordings of Speech) corpus. Features ~100+ diverse speakers with LLM-generated prompts.
Quick Facts
Property
Value
Samples
16,377
Speakers
~100+
Language
English
Format
WebDataset (.tar shards)
Shards
33 (500 samples each)
Source
EARS corpus
Prompt Generation
Gemma-3-4B-IT (creative, nuanced prompts)
Text… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/ears-dramabox.dramabox-tuning-data
DramaBox Tuning Data
Paired training dataset with 12,330 samples designed for fine-tuning DramaBox with bidirectional audio pairs. Combines emotional speech (Emolia) and podcast data in a compact two-part format.
Quick Facts
Property
Value
Samples
12,330 total
Emolia subset
2,316 samples (emotional speech pairs)
Podcast subset
10,014 samples (diverse speaker pairs)
Languages
English, German, Spanish, French
Format
WebDataset (.tar shards)… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/dramabox-tuning-data.echo-robot-dramabox
Echo-TTS Robot — DramaBox-annotated (WER=0)
Robot-voice speech generated with Echo-TTS (EchoDiT preview, jordand/echo-tts-base)
using the Independent sampler preset (cfg_mode=independent, CFG=2,
speaker KV-scale=2), conditioned on a robot reference voice. Across the 40
EmoNet emotion categories, 5 utterances/emotion (200) were written by Gemini,
each generated with 5 seeds (1000 candidates), transcribed with
Parakeet-TDT-0.6B-v3, silence-aware trimmed (500 ms grace, cut in… See the full description on the dataset page: https://huggingface.co/datasets/ChristophSchuhmann/echo-robot-dramabox.podcast-dramabox-dacvae-pairs
podcast-dramabox-dacvae-pairs
Paired audio-codec latents for training a DramaBox → DACVAE latent translator.
Both codecs share an identical grid: 25 Hz, 128-dim, frame-aligned (same length).
Derived from TTS-AGI/podcast-tokenized-bg3.5-enj5.
How it was built (per sample)
DACVAE latent (from source dataset, = target) → DACVAE.decode → 48 kHz mono wav
→ duplicate to stereo → DramaBox/LTX-2.3 audio VAE encode → patchify → DramaBox latent (= input).
Both latents… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/podcast-dramabox-dacvae-pairs.dramabox-gemini-finetune
DramaBox Gemini Fine-Tune Data
Pre-computed DramaBox latent-space training data with speaker embeddings for
fine-tuning with AdaLN-Zero speaker conditioning. All samples are same-speaker
part pairs (reference + target) with Gemini 3.5 Flash prompts.
Stats: 9,397 samples, 94 shards, 5.3 GB
Sample Structure
Each sample in the WebDataset tars contains:
File
Description
{key}_tgt.mp3
Target audio (MP3)
{key}_ref.mp3
Reference speaker audio (MP3)… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/dramabox-gemini-finetune.elise-v2-dramabox-hq-captioned
