datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cocoon-glossesEmotion-Voice-Attribute-Reference-Snippets-DACVAE-Wave
Emotion and Voice Attribute Reference Snippets - DACVAE and Wave
Merged dataset combining TTS-AGI/enhanced-emo-snippets-balanced-DACVAE and
TTS-AGI/emotion-attribute-conditioning-dacvae with decoded WAV audio.
Overview
Total samples: 606,178
Filtered out: 363,331 (samples with speech_quality < 1.8)
Total tar files: 328
Total size: 1.54 TB
Audio format: WAV, 48kHz, PCM 16-bit mono
Latents: DAC-VAE float16 [T, 128] at 25 frames/sec
Dimensions: 57 (40 emotions + 15 voice… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/Emotion-Voice-Attribute-Reference-Snippets-DACVAE-Wave.lrs3_all_video_wavhard_math_wavmimi_wavstn-wavswav2phone-datasetjvnv-nonvデータセットの話者f2で学習したVoiceSpeechMakerのモデルで、jsutコーパスの書き下し分を推論し、推論時のフルコンテキストラベルとペアにしたデータセットです
フルコンテキストラベルは次のような形式です(hfのプレビューは壊れています)
xx^xx-sil+m=i/A:xx+xx+xx/B:xx-xx_xx/C:xx_xx+xx/D:02+xx_xx/E:xx_xx!xx_xx-xx/F:xx_xx#xx_xx@xx_xx|xx_xx/G:3_3%0_0_xx/H:xx_xx/I:xx-xx@xx+xx&xx-xx|xx+xx/J:3_23/K:1+3-23
xx^sil-m+i=z/A:-2+1+3/B:xx-xx_xx/C:02_xx+xx/D:13+xx_xx/E:xx_xx!xx_xx-xx/F:3_3#0_0@1_3|1_23/G:7_2%0_0_1/H:xx_xx/I:3-23@1+1&1-3|1+23/J:xx_xx/K:1+3-23… See the full description on the dataset page: https://huggingface.co/datasets/WariHima/wav2phone-dataset.rxt_10_22_wavicassp_audioWavAgent-GeneralWavAgent-Muti-Generalssl_10_22_wav
