datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sonniss_game_effectsFrench_game_voice
French Game Voice Dataset
Dataset of 100k+ cleaned audio samples of French video game voices with transcriptions.
Features
Format: Mono WAV 16-bit, 48 kHz
Size: ~500 hours of audio
Transcriptions: Faster Whisper Large V3 + in-game subtitles
Usage
from datasets import load_dataset
dataset = load_dataset("Lucari053/French_game_voice")
# Access the data
sample = dataset['train'][0]
audio = sample['audio']
text = sample['text']
Gamehonkai_impact_3rd_game_playthrough
Game Playthrough
最终解析出的语料在 honkai_impact_3rd_chinese_dialogue_corpus。
See honkai_impact_3rd_chinese_dialogue_corpus for final parsed result!
Description (English)
This is a collection of playthrough videos of Honkai Impact 3rd from Hoyoverse, along with efforts to build a Chinese text corpus (with OCR and MLLM-based parsing).
The language setting is Chinese.
All credits to the source author from BiliBili
The dataset contains the following contents:
Videos: The… See the full description on the dataset page: https://huggingface.co/datasets/mrzjy/honkai_impact_3rd_game_playthrough.gametime-outputs
Gametime Outputs
Model outputs (stereo full-duplex mix of user prompt + model response) for the Gametime benchmark.
Load
from datasets import load_dataset
ds = load_dataset("gametime-benchmark/gametime-outputs", "moshi", split="basic")
ex = next(iter(ds))
wav = ex["audio"]["array"] # numpy float array, shape=(n, 2) stereo
sr = ex["audio"]["sampling_rate"] # int (24000 for most, 48000 for gpt-realtime)
print(ex["id"], sr, wav.shape, ex["dataset"])… See the full description on the dataset page: https://huggingface.co/datasets/gametime-benchmark/gametime-outputs.gametime
Gametime Benchmark
The Gametime dataset provides lightweight, streaming-friendly splits for TTS/ASR/SpokenLM prototyping.For full details, please refer to the paper:👉 Game-Time: Evaluating Temporal Dynamics in Spoken Language Models
📦 Download Options
1️⃣ Recommended — Full ZIP Download
If you prefer the original folder layout you can download one of the ZIPs packaged in gametime/download/. There are two kinds available in this repository:… See the full description on the dataset page: https://huggingface.co/datasets/gametime-benchmark/gametime.Diogenes_Gameplay_raw_sample_v01
Diogenes Gameplay Raw Sample v01
Formerly DiogenesLab/Diogenes_COD_sample_v01 — old links redirect here.
A sample dataset. PC gameplay recordings with frame-aligned keyboard/mouse action
annotations, in two batches:
batch
recorded
video
audio
batch 1
2026-07-25/26
1920×1080 @ 30 fps, H.264
none
batch 2
2026-07-31
1920×1080 @ 60 fps, H.264
process-loopback system audio, 48 kHz stereo s16le (zstd-compressed PCM)
This is a sample — the recording output available… See the full description on the dataset page: https://huggingface.co/datasets/DiogenesLab/Diogenes_Gameplay_raw_sample_v01.dsp-gameschool-games
Updaown's UBG
Updaown's UBG is a web proxy with a Clean and Sleek UI and easy-to-use menus. Our goal is to provide the best user experience to everyone.
[!IMPORTANT]
If you fork this project, consider giving it a star in the original repository!
Join Our Discord Community for support, more links, and an active community!
Features
About:Blank Cloaking
Tab Cloaking
Wide collection of apps & games
Clean, Easy-to-use UI
Inspect Element
Various Themes
Password… See the full description on the dataset page: https://huggingface.co/datasets/updaown/school-games.seul-game-processed-30s-razmetkaLEANDRINHO_GAMEPLAYS1LEANDRINHO_GAMEPLAYS2GamesGameCommandLEANDRO_GAMEPLAY1Diogenes_Gameplay_multigame_sample_v01
Diogenes Gameplay Multigame Sample v01
Human gameplay recordings across 19 PC titles, captured with the same
recording pipeline and the same per-frame action-annotation schema on every title.
The point of this sample is breadth: one session per title, identical schema,
per-title semantic action vocabularies (mappings/), and honest per-segment QC
columns computed from the shipped data itself.
19 sessions · 603 segments · 1.72 h · 205,091 annotated frames · 332,572 raw input… See the full description on the dataset page: https://huggingface.co/datasets/DiogenesLab/Diogenes_Gameplay_multigame_sample_v01.ai-arcade-gamesambrosioGames_EduuuLSW1_The_Video_Gamegamecam.mp3Cariocagame-voiceAlex-PatriotaS2T_Splited_Gameshowgameshow_splitgame-novel-xcodec2-example
