CoolFace
Datasetpublic

korovsky/witchspeech

WitchSpeech: Russian voice lines Witcher 3 TTS dataset This is a repack of the dataset so-vits-svc-4.0-ru-The_Witcher_3_Wild_Hunt for TTS purposes.Unlike the original dataset, this dataset also contains transcriptions for voice lines.Transcriptions include stresses for words, even for words with just a single vowel.Additionally, there is a metadata_source.csv file, that contains voice lines text “as-is”. Dataset info (see details in stats.txt): Sample rate: 48 000Total time:… See the full description on the dataset page: https://huggingface.co/datasets/korovsky/witchspeech.

sourceHugging Faceupdated 2y agoView on Hugging Face
5likes22downloads
Dataset Card

WitchSpeech: Russian voice lines Witcher 3 TTS dataset

This is a repack of the dataset so-vits-svc-4.0-ru-The_Witcher_3_Wild_Hunt for TTS purposes. Unlike the original dataset, this dataset also contains transcriptions for voice lines. Transcriptions include stresses for words, even for words with just a single vowel. Additionally, there is a metadata_source.csv file, that contains voice lines text “as-is”.

Dataset info (see details in stats.txt):

Sample rate: 48 000 Total time: 21.49 Number of speakers: 34

In addition to the text, there are also non-speech sound tags in metadata.csv (see non_speech_tags.txt for details). How metadata.csv was created:

  1. 1.Russian voice lines texts were extracted from the game and matched with audio samples from so-vits-svc-4.0-ru-The_Witcher_3_Wild_Hunt dataset.
  2. 2.Stresses for words were automatically added with RUAccent library.
  3. 3.Non-speech sound tags were added from the original voice lines text.

metadata.csv structure: path to audio | text | speaker id (see speaker_ids.json file for details)

This dataset was made possible by the tools and resources created by the following people: Rootreck Den4ikA