CoolFace
Datasetpublic

xincan/Llama-VITS_data

Dataset Card for Llama-VITS_data The dataset repository contains data related with our work "Llama-VITS: Enhancing TTS Synthesis with Semantic Awareness", encapsulating: Filtered dataset EmoV_DB_bea_sem Filelists with semantic embeddings Model checkpoints Human evaluation templates Dataset Details Paper: Llama-VITS: Enhancing TTS Synthesis with Semantic Awareness Curated by: Xincan Feng, Akifumi Yoshimoto Funded by: CyberAgent Inc Repository:… See the full description on the dataset page: https://huggingface.co/datasets/xincan/Llama-VITS_data.

sourceHugging Facemitupdated 2y agoView on Hugging Face
2likes4.7kdownloads
Dataset Card

Dataset Card for Llama-VITS_data

The dataset repository contains data related with our work "Llama-VITS: Enhancing TTS Synthesis with Semantic Awareness", encapsulating:

  • Filtered dataset EmoV_DB_bea_sem
  • Filelists with semantic embeddings
  • Model checkpoints
  • Human evaluation templates

Dataset Details

  • Paper: Llama-VITS: Enhancing TTS Synthesis with Semantic Awareness
  • Curated by: Xincan Feng, Akifumi Yoshimoto
  • Funded by: CyberAgent Inc
  • Repository: https://github.com/xincanfeng/vitsGPT
  • Demo: https://xincanfeng.github.io/Llama-VITS_demo/

Dataset Creation

We fileterd EmoV_DB_bea_sem dataset from EmoV_DB (Adigwe et al., 2018), a database of emotional speech containing data for male and female actors in English and French. EmoVDB covers 5 emotion classes, amused, angry, disgusted, neutral, and sleepy. To factor out the effect of different speakers, we filtered the original EmoVDB dataset into the speech of a specific female English speaker, bea. Then we use Llama2 to predict the emotion label of the transcript chosen from the above 5 emotion classes, and select the audio samples which has the same predicted emotion. The filtered dataset contains 22.8-minute records for training. We named the filtered dataset EmoV_DB_bea_sem and investigated how the semantic embeddings from Llama2 behave in naturalness and expressiveness on it. Please refer to our paper for more information.

Citation

If our work is useful to you, please cite our paper: "Llama-VITS: Enhancing TTS Synthesis with Semantic Awareness".

sh
@misc{feng2024llamavits,
      title={Llama-VITS: Enhancing TTS Synthesis with Semantic Awareness}, 
      author={Xincan Feng and Akifumi Yoshimoto},
      year={2024},
      eprint={2404.06714},
      archivePrefix={arXiv},
      primaryClass={cs.CL}
}