SynDataLab-JA-Refs/Irodori-Ja-Spk3-10k
SynDataLab/Irodori-Ja-Spk3-10k 10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker family — see also Spk1, Spk2, Spk3, Spk4 repos. Speaker (Spk3): 40s male, low calm mature — 40代男性、低めで穏やかな落ち着いた声. How this speaker was made The voice identity for Spk3 was created in two stages: Stage 1 — voice anchor design. Using Irodori-TTS-500M-v2-VoiceDesign, several candidate audio samples were synthesized from the… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/Irodori-Ja-Spk3-10k.
SynDataLab/Irodori-Ja-Spk3-10k
10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker family — see also Spk1, Spk2, Spk3, Spk4 repos.
Speaker (Spk3): 40s male, low calm mature — 40代男性、低めで穏やかな落ち着いた声.
How this speaker was made
The voice identity for Spk3 was created in two stages:
Stage 1 — voice anchor design. Using Irodori-TTS-500M-v2-VoiceDesign, several candidate audio samples were synthesized from the Japanese caption below and varied seeds. Then one was hand-picked as the canonical anchor for this speaker:
`` 40代男性の低めで穏やかな声。落ち着いた雰囲気で、ゆったりとした会話調。声優風ではなく、自然な大人の男性の話し方。 ``Stage 2 — bulk synthesis with anchor as reference. The anchor wav was passed as ref_wav to the base Irodori-TTS-500M-v2 model for all 10,000 generations. The base model's speaker-conditioning enforces identity consistency across the run — every clip in this repo is the same voice.
How the texts were made
10,000 unique conversational Japanese single utterances were generated via the DeepSeek Pro API (deepseek-v4-pro, thinking mode explicitly disabled). The prompting strategy:
- Anchored each batch on one of 10 emotional buckets (joy, sadness, anger, surprise, embarrassment, nostalgia, love, exhaustion, curiosity, fear) × concrete situations.
- Three length buckets (quip 12–35, mid 35–75, long 75–130 full-width chars), biased toward
midfor the 5–9 s spoken sweet spot. - Filtered out: bracketed paralinguistic tags (
[笑い],[ため息], etc.), English letters, AI-disclaimer phrases, therapist-cliché phrases, banned openers. - Globally deduplicated (NFKC-normalized), then deterministically shuffled (seed 42) and split into 4 disjoint partitions of 10,000 lines — one per speaker. *No text appears in more than one Spk repo.**
Schema
Stats
- 10,000 unique (text, audio) pairs
- Mean duration ≈ 7.5–8.3 s, P99 ≈ 23–27 s
- 0 truncated clips (<1.5 s), 0 generation failures
- Total audio: ~21 hours
Synthesis params
- model:
Aratako/Irodori-TTS-500M-v2 - numsteps: 40, cfgscaletext: 3.0, cfgscale_speaker: 5.0
- bf16 inference on H100
- seed: deterministic per row (= row index)
- reference:
Spk3anchor wav, codec-encoded once and reused
License
CC-BY-NC-SA-4.0 — non-commercial, share-alike. Inherits from the Irodori-TTS-500M-v2 model weights license.
