mcamara/all-words-in-english-with-pink-trombone
Dataset Card for Pink Trombone English Phonetic & Landmark Dataset Repository: mcamara/all-words-in-english-with-pink-trombone Modality: Audio + time-aligned events (landmarks) + articulatory keyframes Language: English (IPA) Sampling rate: 44,100 Hz (mono) Voices: two synthetic voices — M (male) and F (female) Summary A large-scale, clean synthetic speech dataset generated with the Pink Trombone articulatory synthesizer. Every English dictionary word is… See the full description on the dataset page: https://huggingface.co/datasets/mcamara/all-words-in-english-with-pink-trombone.
Dataset Card for Pink Trombone English Phonetic & Landmark Dataset
Repository: mcamara/all-words-in-english-with-pink-trombone Modality: Audio + time-aligned events (landmarks) + articulatory keyframes Language: English (IPA) Sampling rate: 44,100 Hz (mono) Voices: two synthetic voices — M (male) and F (female)
Summary
A large-scale, clean synthetic speech dataset generated with the Pink Trombone articulatory synthesizer. Every English dictionary word is synthesized in two voices (male and female). Each example links a word to:
- audio — the synthesized waveform (44.1 kHz, mono, FLAC — lossless),
- utterance — the articulatory keyframes driving the synthesis (
{name, keyframes}), - landmarks — time-aligned acoustic landmarks detected from the audio.
Fields
Landmark types
Voice parameters
Generation
Produced with the Pink Trombone web synthesizer driven headlessly via Playwright (pink-trombone-demos/batch-generator). For each word: text → IPA → articulatory keyframes (TTS module) → real-time synthesis + recording (Pink Trombone module) → WAV + landmark extraction (LEXI module). Landmarks are detected from energy/articulatory events. Audio is the synthesizer's native 44.1 kHz output (no resampling), stored as lossless FLAC.
Loading
from datasets import load_dataset
ds = load_dataset("mcamara/all-words-in-english-with-pink-trombone", split="train")
ex = ds[0]
ex["audio"] # {'array': ..., 'sampling_rate': 44100}
ex["id"], ex["sex"]
import json
json.loads(ex["landmarks"])
json.loads(ex["utterance"])
# filter one voice
male = ds.filter(lambda r: r["sex"] == "M")License
MIT.
