SynDataLab-JA-Refs/irodori-refs-10k-v2
Irodori TTS Reference Voices v2 (10K) 10,000 reference voices generated with Aratako/Irodori-TTS-500M-v2-VoiceDesign (no_ref=True) using a richer caption space than v1: 8 axes (gender × age × pitch × tone × speed × distance × emotion × quality) with per-voice unique caption combinations, plus gender alternation, an incompatibility filter (no contradictory "whisper + speak loudly" combos), and a similarity-rejection window so consecutive voices stay distinct. Each ref's text… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/irodori-refs-10k-v2.
Irodori TTS Reference Voices v2 (10K)
10,000 reference voices generated with Aratako/Irodori-TTS-500M-v2-VoiceDesign (no_ref=True) using a richer caption space than v1: 8 axes (gender × age × pitch × tone × speed × distance × emotion × quality) with per-voice unique caption combinations, plus gender alternation, an incompatibility filter (no contradictory "whisper + speak loudly" combos), and a similarity-rejection window so consecutive voices stay distinct.
Each ref's text comes from the v2 corpus (SynDataLab/DeepSeekFlash-3M-en-style conversational utterances, Japanese) — 1/3 from each of the quip / mid / long length buckets so audio spans roughly 4-15 s.
Schema
Related
Companion clones dataset: SynDataLab/irodori-clones-3m-v2 (300 clones per voice → 3M total ≈ 7000 h of audio).
