CoolFace
Datasetpublic

SynDataLab-JA-Refs/irodori-refs-10k-v2

Irodori TTS Reference Voices v2 (10K) 10,000 reference voices generated with Aratako/Irodori-TTS-500M-v2-VoiceDesign (no_ref=True) using a richer caption space than v1: 8 axes (gender × age × pitch × tone × speed × distance × emotion × quality) with per-voice unique caption combinations, plus gender alternation, an incompatibility filter (no contradictory "whisper + speak loudly" combos), and a similarity-rejection window so consecutive voices stay distinct. Each ref's text… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/irodori-refs-10k-v2.

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes56downloads
Dataset Card

Irodori TTS Reference Voices v2 (10K)

10,000 reference voices generated with Aratako/Irodori-TTS-500M-v2-VoiceDesign (no_ref=True) using a richer caption space than v1: 8 axes (gender × age × pitch × tone × speed × distance × emotion × quality) with per-voice unique caption combinations, plus gender alternation, an incompatibility filter (no contradictory "whisper + speak loudly" combos), and a similarity-rejection window so consecutive voices stay distinct.

Each ref's text comes from the v2 corpus (SynDataLab/DeepSeekFlash-3M-en-style conversational utterances, Japanese) — 1/3 from each of the quip / mid / long length buckets so audio spans roughly 4-15 s.

Schema

columntype
audioaudio(sampling_rate=48000), mono
textstring (the spoken text, includes emoji prosody cues)
speaker_idref_speaker_v2_NNNNN, sequential 1-10000

Related

Companion clones dataset: SynDataLab/irodori-clones-3m-v2 (300 clones per voice → 3M total ≈ 7000 h of audio).