datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wuwa-voice-EN
wuwa-voice-EN
wuwa voice EN is a dataset of voice line from Wuthering Waves
Attribute
Value
Language
English
Total Samples
37,805
Total Duration
~24 GB (WAV format)
Unique Speakers
913
Categories
106
Audio Format
WAV
Transcription Format
Plain text
Field
Description
file_name
Relative path to the audio file (e.g., data/en_vo_Category_1_1.wav)
text
transcript
speaker
Character (e.g., Zani, Carlotta, {PlayerName})
speaker_id
Numeric ID… See the full description on the dataset page: https://huggingface.co/datasets/igidn/wuwa-voice-EN.WuwaWuWa_Mapsaviationqa121huatuo_encyclopedia_qa
Dataset Card for Huatuo_encyclopedia_qa
Dataset Summary
This dataset has a total of 364,420 pieces of medical QA data, some of which have multiple questions in different ways. We extract medical QA pairs from plain texts (e.g., medical encyclopedias and medical articles). We collected 8,699 encyclopedia entries for diseases and 2,736 encyclopedia entries for medicines on Chinese Wikipedia. Moreover, we crawled 226,432 high-quality medical articles from the Qianwen… See the full description on the dataset page: https://huggingface.co/datasets/wuwu616/huatuo_encyclopedia_qa.huatuo_knowledge_graph_qa
Dataset Card for Huatuo_knowledge_graph_qa
Dataset Summary
We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map.
Dataset Creation
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/wuwu616/huatuo_knowledge_graph_qa.tiny_cocooutputpvf_benchmark_backupprivate-serverWuWakids-scifi-writer-dataset
原始数据集
writingprompts
XueyingJia/Children-Stories-Collection-Filtered
处理
从上述数据集中提取出科幻故事数据集、儿童故事数据集、儿童科幻故事数据集。
三个数据集system_prompt各不相同,三个数据集打乱合并成最后数据集
LAFAN1_Retargeting_Dataset
LAFAN1 Retargeting Dataset
To make the motion of humanoid robots more natural, we retargeted LAFAN1 motion capture data to Unitree's humanoid robots, supporting three models: H1, H1_2, and G1. This retargeting was achieved through numerical optimization based on Interaction Mesh and IK, considering end-effector pose constraints, as well as joint position and velocity constraints, to prevent foot slippage. It is important to note that the retargeting only accounted for kinematic… See the full description on the dataset page: https://huggingface.co/datasets/WUWWUWWUW/LAFAN1_Retargeting_Dataset.wuw_dataset
Yougen/wuw_dataset
Wake-Up-Word (WUW) speech dataset, packed as WebDataset tar shards.
The input is a Kaldi-style data directory
(wav.scp, text, utt2spk, utt2dur, segments), where multiple
utterances share a long recording via the segments file.
To avoid duplicating audio, each tar sample corresponds to one full
recording. The utterance-level metadata (id / start / end / text / spk / duration) is stored in a JSON list inside that sample.
Downstream consumers slice the decoded… See the full description on the dataset page: https://huggingface.co/datasets/bhyuan/wuw_dataset.wuw_min
ygyuan/wuw_min
Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards.
The input is a Kaldi-style data directory
(wav.scp, text, utt2spk, utt2dur, segments), where each
utterance is packed as a single tar sample.
Layout
data/
<split>/
metadata.csv
audio/
<split>-000.tar
<split>-001.tar
...
Shard counts:
train: 65 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
<key>.wav # raw audio… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/wuw_min.wuw_cszhen
ygyuan/wuw_cszhen
Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards.
The input is a Kaldi-style data directory
(wav.scp, text, utt2spk, utt2dur, segments), where each
utterance is packed as a single tar sample.
Layout
data/
<split>/
metadata.csv
audio/
<split>-000.tar
<split>-001.tar
...
Shard counts:
train: 712 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
<key>.wav # raw audio… See the full description on the dataset page: https://huggingface.co/datasets/heimayuan/wuw_cszhen.wuwa_dialogue
Dialogues Dataset for translator LLM (Wuthering Waves ONLY)
Project Readme 👇
แปล Dialogue เกมเป็นภาษาไทยด้วย LLM
Report Issues
About The Project 😃
บางครั้งเรามีเกมที่เราชอบมากๆ แต่น่าเสียดายที่เกมนั้นไม่มีแปลไทย (โคตร Sad) จะนั่งแปลเองก็เหนื่อยมากๆ ยิ่งเป็นเกมเนื้อเรื่องยาวๆ 6-8ชม. นี่ไม่ไหวเลย เราก็เลยคิดเอา LLM มาช่วยแปล ตอนแรกก็ว่าจะเก็บไว้ใช้คนเดียว เพราะ CodeBase… See the full description on the dataset page: https://huggingface.co/datasets/kang49/wuwa_dialogue.dmwuw_accent
ygyuan/wuw_accent
Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards.
The input is a Kaldi-style data directory
(wav.scp, text, utt2spk, utt2dur, segments), where each
utterance is packed as a single tar sample.
Layout
data/
<split>/
metadata.csv
audio/
<split>-000.tar
<split>-001.tar
...
Shard counts:
train: 509 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
<key>.wav # raw audio… See the full description on the dataset page: https://huggingface.co/datasets/Yougen/wuw_accent.IML_comparesppy36-spwuw_wuy
ygyuan/wuw_wuy
Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards.
The input is a Kaldi-style data directory
(wav.scp, text, utt2spk, utt2dur, segments), where each
utterance is packed as a single tar sample.
Layout
data/
<split>/
metadata.csv
audio/
<split>-000.tar
<split>-001.tar
...
Shard counts:
train: 99 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
<key>.wav # raw audio… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/wuw_wuy.wuw_testset1
Yougen/wuw_testset1
Wake-Up-Word (WUW) speech dataset, packed as WebDataset tar shards.
The input is a Kaldi-style data directory
(wav.scp, text, utt2spk, utt2dur, segments), where multiple
utterances share a long recording via the segments file.
To avoid duplicating audio, each tar sample corresponds to one full
recording. The utterance-level metadata (id / start / end / text / spk / duration) is stored in a JSON list inside that sample.
Downstream consumers slice the decoded… See the full description on the dataset page: https://huggingface.co/datasets/heimayuan/wuw_testset1.chunWuWa_Frame_GenTitanbrainv1PediatricIMLDatasetwuweibu
