CoolFace
Datasetpublic

igidn/wuwa-voice-EN

wuwa-voice-EN wuwa voice EN is a dataset of voice line from Wuthering Waves Attribute Value Language English Total Samples 37,805 Total Duration ~24 GB (WAV format) Unique Speakers 913 Categories 106 Audio Format WAV Transcription Format Plain text Field Description file_name Relative path to the audio file (e.g., data/en_vo_Category_1_1.wav) text transcript speaker Character (e.g., Zani, Carlotta, {PlayerName}) speaker_id Numeric ID… See the full description on the dataset page: https://huggingface.co/datasets/igidn/wuwa-voice-EN.

sourceHugging Faceupdated 4mo agoView on Hugging Face
2likes586downloads
Dataset Card

wuwa-voice-EN

wuwa voice EN is a dataset of voice line from Wuthering Waves

AttributeValue
LanguageEnglish
Total Samples37,805
Total Duration~24 GB (WAV format)
Unique Speakers913
Categories106
Audio FormatWAV
Transcription FormatPlain text
FieldDescription
file_nameRelative path to the audio file (e.g., data/en_vo_Category_1_1.wav)
texttranscript
speakerCharacter (e.g., Zani, Carlotta, {PlayerName})
speaker_idNumeric ID assigned to the speaker
audio_keyInternal game audio identifier
gender_suffixF or M for player-character variants; null otherwise
categoryContent category (story arc, character quest, event, etc.)
confidenceTranscription confidence: high, low, or auto
plotaudio_settingIn-game audio context (e.g., subtitle_normal)

All annotations are stored in metadata.jsonl, where each line is a JSON object corresponding to one audio file.

Confidence

  • High: 35,273 samples (93.3%) — manually verified or high-confidence matches.
  • Low: 1,277 samples (3.4%) — manually flagged as uncertain or partial matches.
  • Auto: 1,255 samples (3.3%) — the transcripts automatically generated by OpenAI Whisper; not manually reviewed.

Player Character Gender Variants

player-character lines have separate F (female) and M (male):

  • Female (`F`): 3,901 samples
  • Male (`M`): 3,901 samples

Licensing information

Copyright © kurogames. All Rights Reserved.