datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GTSinger
GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks
Yu Zhang*, Changhao Pan*, Wenxiang Guo*, Ruiqi Li, Zhiyuan Zhu, Jialei Wang, Wenhao Xu, Jingyu Lu, Zhiqing Hong, Chuxin Wang, LiChao Zhang, Jinzheng He, Ziyue Jiang, Yuxin Chen, Chen Yang, Jiecheng Zhou, Xinyu Cheng, Zhou Zhao | Zhejiang University
Dataset of GTSinger (NeurIPS 2024 Spotlight): A Global Multi-Technique Singing Corpus with Realistic Music Scores for All… See the full description on the dataset page: https://huggingface.co/datasets/AaronZ345/GTSinger.Neko_Audio-30K_LongOpenGameArt-CC-BY-SA-3.0
Dataset Card for OpenGameArt-CC-BY-SA-3.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the Creative Commons Attribution-ShareAlike 3.0 Unported (CC-BY-SA-3.0) license. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata.
Languages
The dataset is primarily… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-CC-BY-SA-3.0.OpenGameArt-OGA-BY-3.0
Dataset Card for OpenGameArt-OGA-BY-3.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the OpenGameArt Attribution (OGA-BY-3.0) license. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and associated metadata.
Languages
The dataset is primarily monolingual:
English (en): All asset descriptions and metadata are in… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-OGA-BY-3.0.OpenGameArt-CC-BY-3.0
Dataset Card for OpenGameArt-CC-BY-3.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the Creative Commons Attribution 3.0 (CC-BY-3.0) license. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata.
Languages
The dataset is primarily monolingual:
English (en): All… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-CC-BY-3.0.Seamless_Dummy_Dataset_Fixed_3
MMLU-Pro json
This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details.
odubench
ODU-Bench: Omni Demand Understanding
Introduction
ODU-Bench is a benchmark for contextual user-intent inference in audio
and audio-visual interaction. Omni Demand Understanding (ODU) asks a model to
detect whether a valid user demand is present and infer the user's intent from
multimodal and conversational context. A demand is the outcome a user wants an
assistant to achieve, including response-relevant objects, constraints and
trigger conditions.
The… See the full description on the dataset page: https://huggingface.co/datasets/qc316/odubench.H3-Character-Swap-v1
H3 Character Swap v1
A reference-conditioned character-replacement dataset for MiniMax H3 Ref2VA LoRA training with Ostris AI Toolkit. It combines synthetic still-image edits with unchanged real-motion regularization videos.
134 examples: 94 character-swap edits and 40 preservation clips. Training has 76 edits + 32 clips; validation has 18 edits + 8 clips. Prepared resolution is 1344×768 at 24 fps. The companion 1,000-step LoRA are available separately.
Task and… See the full description on the dataset page: https://huggingface.co/datasets/akatz-ai/H3-Character-Swap-v1.EGYSpeak
EGYSpeak
A curated dataset of 147,979 single-speaker Egyptian Arabic (pure dialect) audio clips with transcriptions, sourced from the fadisarwat/egyptian-arabic-lines Kaggle dataset and processed through a rigorous ASR pipeline.
Quick Start
1. Download the dataset:
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="MohamedGomaa30/EGYSpeak",
repo_type="dataset",
local_dir="EGYSpeak",
)
2. Extract the dataset:
from… See the full description on the dataset page: https://huggingface.co/datasets/MohamedGomaa30/EGYSpeak.OpenGameArt-GPL-3.0
Dataset Card for OpenGameArt-GPL-3.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the GNU General Public License version 3.0 (GPL-3.0). The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, and textures along with their associated metadata.
Languages
The dataset is primarily monolingual:
English (en): All asset descriptions… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-GPL-3.0.OpenGameArt-GPL-3.0
Dataset Card for OpenGameArt-GPL-3.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the GNU General Public License version 3.0 (GPL-3.0). The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, and textures along with their associated metadata.
Languages
The dataset is primarily monolingual:
English (en): All asset descriptions… See the full description on the dataset page: https://huggingface.co/datasets/irfankabir02/OpenGameArt-GPL-3.0.test321
test321
This is a merged speech dataset containing 118 audio segments from 2 source datasets.
Dataset Information
Total Segments: 118
Speakers: 4
Languages: tr
Emotions: happy, angry, sad, neutral
Original Datasets: 2
Dataset Structure
Each example contains:
audio: Audio file (WAV format, 16kHz sampling rate)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged datasets)
emotion: Detected emotion… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/test321.vggsync-3k
VGGSync-3K · out-of-domain audio-visual sync benchmark
Out-of-domain evaluation set used in the paper
When Vision Speaks for Sound.
Derived from VGGSoundSync,
this 3,000-clip slice tests whether a video-capable MLLM can detect
audio temporal offsets on everyday sound events outside the THUD
in-domain training distribution.
Each item is one VGGSound clip in one of three conditions:
Condition
Count
gt_synced
gt_direction
gt_offset_sec
Audio aligned (no shift)
1,000
true… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/vggsync-3k.qwen3-tts-preset-voices
Qwen3-TTS preset voice embeddings
The 9 named speakers from Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice packaged as a small sidecar bundle usable with the -Base checkpoint.
bundle.safetensors — 9 × 2048-d bfloat16 rows, ~37 KB total
bundle.json — metadata (speaker name → spk_id, gender, supported languages)
Each row is lifted from talker.model.codec_embedding.weight in the CustomVoice checkpoint at the speaker-ID index from its config.json. With these rows, you can:
Deploy only… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qwen3-tts-preset-voices.yt3_chunked_tokenized
yt3_chunked_48k_tokenized
This is a gated Russian tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: ru (Russian)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt3_chunked_tokenized.whisper-finetune-audio_test3chester_bennington_30tracks_datasetaitf-dfk3-synthetic-audio-datasetlntest3
