datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
starrail-voice
StarRail Voice
StarRail Voice is a dataset of voice lines from the popular game Honkai: Star Rail.
Hugging Face 🤗 StarRail-Voice
ModelScope StarRail-Voice
Last update at 2026-07-16, game version 4.4.0
403437 wavs
60164 without speaker (15%)
61375 without transcription (15%)
57869 without inGameFilename (14%)
Dataset Details
Dataset Description
The dataset contains voice lines from the game's characters in multiple languages, including Chinese… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/starrail-voice.Kurisu_Voice
Makise Kurisu Multilingual Voice Dataset
13,999 labelled clips (17.1423 hours) of Makise Kurisu,
including the Amadeus Kurisu variant, in six languages, cut from the
STEINS;GATE games, the anime, and three character songs.
Every clip carries the transcript, a measured acoustic profile, an independent
speaker-identity check, and — where it could be earned rather than guessed — an
expressive tag. This is an unofficial, fan-made dataset with no affiliation to
the STEINS;GATE rights… See the full description on the dataset page: https://huggingface.co/datasets/starrydark/Kurisu_Voice.honkai-star-rail-voices
Honkai: Star Rail — Voice Lines (Multi-Language)
An archive of character voice data extracted from Honkai: Star Rail (崩壊:スターレイル), repackaged as Parquet shards per audio language.
Dataset Summary
Field
Value
Game
Honkai: Star Rail (崩壊:スターレイル)
Publisher
HoYoverse / miHoYo Co., Ltd.
Languages
中文 (zh), 日本語 (ja), English (en), 한국어 (ko)
Game version
3.8
Source format
WAV + sidecar transcripts (.lab / .txt)
Distribution format
Apache Parquet (zstd), ~500 MiB… See the full description on the dataset page: https://huggingface.co/datasets/ultemica/honkai-star-rail-voices.ernest_heminguei_stary_chalavek_i_mora_all
Стары чалавек і мора
Аўтар / Author: Эрнэст ХемінгуэйМова / Language: Беларуская (Belarusian)
Аўдыё нарэзана з арыгінальнага запісу ў зыходнай частаце дыскрэтызацыі (native), мона, фрагменты да 30 секунд.
Частка калекцыі Belarusian Audiobooks (native).
Радкоў у датасеце
1,238
Працягласць
4 гадз 6 хв
Частата дыскрэтызацыі
44100 Hz
Каналы
мона
Даўжыня фрагмента
да 30 с
Структура
Кожны радок змяшчае:
audio — аўдыёфрагмент (native SR, мона… See the full description on the dataset page: https://huggingface.co/datasets/fosters/ernest_heminguei_stary_chalavek_i_mora_all.ernest_heminguei_stary_chalavek_i_mora_output_original
Стары чалавек і мора — арыгінальнае аўдыё
Аўтар / Author: Эрнэст ХемінгуэйМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
ernest_heminguei_stary_chalavek_i_mora_output
Доўгасць аўдыё
4h30m
Радкоў у датасеце
995
Структура
Кожны радок змяшчае:
audio — арыгінальны… See the full description on the dataset page: https://huggingface.co/datasets/fosters/ernest_heminguei_stary_chalavek_i_mora_output_original.
