CoolFace
Datasetpublic

SynDataLab-JA-Refs/Irodori-Ja-Spk3-10k

SynDataLab/Irodori-Ja-Spk3-10k 10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker family — see also Spk1, Spk2, Spk3, Spk4 repos. Speaker (Spk3): 40s male, low calm mature — 40代男性、低めで穏やかな落ち着いた声. How this speaker was made The voice identity for Spk3 was created in two stages: Stage 1 — voice anchor design. Using Irodori-TTS-500M-v2-VoiceDesign, several candidate audio samples were synthesized from the… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/Irodori-Ja-Spk3-10k.

sourceHugging Facecc-by-nc-sa-4.0updated 5mo agoView on Hugging Face
0likes95downloads
Dataset Card

SynDataLab/Irodori-Ja-Spk3-10k

10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker family — see also Spk1, Spk2, Spk3, Spk4 repos.

Speaker (Spk3): 40s male, low calm mature — 40代男性、低めで穏やかな落ち着いた声.

How this speaker was made

The voice identity for Spk3 was created in two stages:

Stage 1 — voice anchor design. Using Irodori-TTS-500M-v2-VoiceDesign, several candidate audio samples were synthesized from the Japanese caption below and varied seeds. Then one was hand-picked as the canonical anchor for this speaker:

`` 40代男性の低めで穏やかな声。落ち着いた雰囲気で、ゆったりとした会話調。声優風ではなく、自然な大人の男性の話し方。 ``

Stage 2 — bulk synthesis with anchor as reference. The anchor wav was passed as ref_wav to the base Irodori-TTS-500M-v2 model for all 10,000 generations. The base model's speaker-conditioning enforces identity consistency across the run — every clip in this repo is the same voice.

How the texts were made

10,000 unique conversational Japanese single utterances were generated via the DeepSeek Pro API (deepseek-v4-pro, thinking mode explicitly disabled). The prompting strategy:

  • Anchored each batch on one of 10 emotional buckets (joy, sadness, anger, surprise, embarrassment, nostalgia, love, exhaustion, curiosity, fear) × concrete situations.
  • Three length buckets (quip 12–35, mid 35–75, long 75–130 full-width chars), biased toward mid for the 5–9 s spoken sweet spot.
  • Filtered out: bracketed paralinguistic tags ([笑い], [ため息], etc.), English letters, AI-disclaimer phrases, therapist-cliché phrases, banned openers.
  • Globally deduplicated (NFKC-normalized), then deterministically shuffled (seed 42) and split into 4 disjoint partitions of 10,000 lines — one per speaker. *No text appears in more than one Spk repo.**

Schema

ColumnTypeDescription
audioAudio(48000)48 kHz mono WAV
textstringJapanese text spoken in the clip
speaker_idstringSpk3 — same value for every row in this repo

Stats

  • 10,000 unique (text, audio) pairs
  • Mean duration ≈ 7.5–8.3 s, P99 ≈ 23–27 s
  • 0 truncated clips (<1.5 s), 0 generation failures
  • Total audio: ~21 hours

Synthesis params

  • model: Aratako/Irodori-TTS-500M-v2
  • numsteps: 40, cfgscaletext: 3.0, cfgscale_speaker: 5.0
  • bf16 inference on H100
  • seed: deterministic per row (= row index)
  • reference: Spk3 anchor wav, codec-encoded once and reused

License

CC-BY-NC-SA-4.0 — non-commercial, share-alike. Inherits from the Irodori-TTS-500M-v2 model weights license.