Galgame
Datasets
All datasets matching “Galgame”Galgame-VisualNovel-Reupload
Galgame VisualNovel Reupload
This repository is a reupload of the visual novel dataset OOPPEENN/56697375616C4E6F76656C5F44617461736574.
The goal of this reupload is to restructure the data for easier and more efficient use with the datasets library, instead of having to manually extract each archive file and parse json files of the original dataset.
Loading the entire dataset
To load and stream all voice lines from all games combined, simply load the train split. The… See the full description on the dataset page: https://huggingface.co/datasets/joujiboi/Galgame-VisualNovel-Reupload.Galgame_Speech_SER_16kHz
Dataset Card for Galgame_Speech_SER_16kHz
[!IMPORTANT]The following rules (in the original repository) must be followed:
必须遵守GNU General Public License v3.0内的所有协议!附加:禁止商用,本数据集以及使用本数据集训练出来的任何模型都不得用于任何商业行为,如要用于商业用途,请找数据列表内的所有厂商授权(笑),因违反开源协议而出现的任何问题都与本人无关!
训练出来的模型必须开源,是否在README内引用本数据集由训练者自主决定,不做强制要求。
English:
You must comply with all the terms of the GNU General Public License v3.0!Additional note: Commercial use is prohibited. This dataset and any model trained using this dataset… See the full description on the dataset page: https://huggingface.co/datasets/litagin/Galgame_Speech_SER_16kHz.Galgame_Speech_ASR_16kHz
Dataset Card for Galgame_Speech_ASR_16kHz
[!IMPORTANT]The following rules (in the original repository) must be followed:
必须遵守GNU General Public License v3.0内的所有协议!附加:禁止商用,本数据集以及使用本数据集训练出来的任何模型都不得用于任何商业行为,如要用于商业用途,请找数据列表内的所有厂商授权(笑),因违反开源协议而出现的任何问题都与本人无关!
训练出来的模型必须开源,是否在README内引用本数据集由训练者自主决定,不做强制要求。
English:
You must comply with all the terms of the GNU General Public License v3.0!Additional note: Commercial use is prohibited. This dataset and any model trained using this dataset… See the full description on the dataset page: https://huggingface.co/datasets/litagin/Galgame_Speech_ASR_16kHz.Galgame_Gemini_Captions
Galgame_Gemini_Captions
Dataset Description
This dataset consists of audio data, their corresponding transcriptions, and detailed audio captions generated by Gemini 2.5 Pro. The data is a subset of the OOPPEENN/56697375616C4E6F76656C5F4461736574 dataset.
It is intended for training Text-to-Speech (TTS) models that can be controlled via descriptive metadata tags (e.g., emotion, speaker profile, style).
Dataset Structure
The dataset is divided into two… See the full description on the dataset page: https://huggingface.co/datasets/NandemoGHS/Galgame_Gemini_Captions.Galgame_Dataset_stats
Viewer in 🤗 spaces here!
Statistics of the audio (voice) files in the OOPPEENN/Galgame_Dataset (in TSV format):
game_name num_speakers num_mono_files num_stereo_files num_error_files total_duration_hours avg_sample_rate_kHz avg_precision avg_bitrate_kbps codec total_size_GB
Game1 47 15055 1 8 20.08 48.0 16.0 88.49 Vorbis 0.73
Game2 40 15370 0 7 30.10 47.8 16.0 87.113 Vorbis 1.07
...
For each game, the specs folder includes spectrogram images from 5 randomly selected audio files.… See the full description on the dataset page: https://huggingface.co/datasets/litagin/Galgame_Dataset_stats.minoricloud-galgame-vocal-mp3-collection-1989-2023
