CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AaronZ345 /GTSinger GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks Yu Zhang*, Changhao Pan*, Wenxiang Guo*, Ruiqi Li, Zhiyuan Zhu, Jialei Wang, Wenhao Xu, Jingyu Lu, Zhiqing Hong, Chuxin Wang, LiChao Zhang, Jinzheng He, Ziyue Jiang, Yuxin Chen, Chen Yang, Jiecheng Zhou, Xinyu Cheng, Zhou Zhao | Zhejiang University Dataset of GTSinger (NeurIPS 2024 Spotlight): A Global Multi-Technique Singing Corpus with Realistic Music Scores for All… See the full description on the dataset page: https://huggingface.co/datasets/AaronZ345/GTSinger.audiotext-to-audio10K<n<100K17 likes22k downloads1y agoHugging Face02liumindmind /Neko_Audio-30K_Longaudio10K<n<100K7 likes3k downloads4mo agoHugging Face03nyuuzyou /OpenGameArt-CC-BY-SA-3.0 Dataset Card for OpenGameArt-CC-BY-SA-3.0 Dataset Summary This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the Creative Commons Attribution-ShareAlike 3.0 Unported (CC-BY-SA-3.0) license. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata. Languages The dataset is primarily… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-CC-BY-SA-3.0.audioimage-classification1K<n<10K1 likes584 downloads1y agoHugging Face04nyuuzyou /OpenGameArt-OGA-BY-3.0 Dataset Card for OpenGameArt-OGA-BY-3.0 Dataset Summary This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the OpenGameArt Attribution (OGA-BY-3.0) license. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and associated metadata. Languages The dataset is primarily monolingual: English (en): All asset descriptions and metadata are in… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-OGA-BY-3.0.audioimage-classificationn<1K1 likes482 downloads1y agoHugging Face05nyuuzyou /OpenGameArt-CC-BY-3.0 Dataset Card for OpenGameArt-CC-BY-3.0 Dataset Summary This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the Creative Commons Attribution 3.0 (CC-BY-3.0) license. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata. Languages The dataset is primarily monolingual: English (en): All… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-CC-BY-3.0.textimage-classification1K<n<10K1 likes387 downloads1y agoHugging Face06rajjanardhan00 /Seamless_Dummy_Dataset_Fixed_3 MMLU-Pro json This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details. audioquestion-answeringn<1K0 likes282 downloads1y agoHugging Face07qc316 /odubench ODU-Bench: Omni Demand Understanding Introduction ODU-Bench is a benchmark for contextual user-intent inference in audio and audio-visual interaction. Omni Demand Understanding (ODU) asks a model to detect whether a valid user demand is present and infer the user's intent from multimodal and conversational context. A demand is the outcome a user wants an assistant to achieve, including response-relevant objects, constraints and trigger conditions. The… See the full description on the dataset page: https://huggingface.co/datasets/qc316/odubench.audio1K<n<10K3 likes254 downloads7d agoHugging Face08akatz-ai /H3-Character-Swap-v1 H3 Character Swap v1 A reference-conditioned character-replacement dataset for MiniMax H3 Ref2VA LoRA training with Ostris AI Toolkit. It combines synthetic still-image edits with unchanged real-motion regularization videos. 134 examples: 94 character-swap edits and 40 preservation clips. Training has 76 edits + 32 clips; validation has 18 edits + 8 clips. Prepared resolution is 1344×768 at 24 fps. The companion 1,000-step LoRA are available separately. Task and… See the full description on the dataset page: https://huggingface.co/datasets/akatz-ai/H3-Character-Swap-v1.imagen<1K0 likes231 downloads1d agoHugging Face09MohamedGomaa30 /EGYSpeak EGYSpeak A curated dataset of 147,979 single-speaker Egyptian Arabic (pure dialect) audio clips with transcriptions, sourced from the fadisarwat/egyptian-arabic-lines Kaggle dataset and processed through a rigorous ASR pipeline. Quick Start 1. Download the dataset: from huggingface_hub import snapshot_download snapshot_download( repo_id="MohamedGomaa30/EGYSpeak", repo_type="dataset", local_dir="EGYSpeak", ) 2. Extract the dataset: from… See the full description on the dataset page: https://huggingface.co/datasets/MohamedGomaa30/EGYSpeak.textautomatic-speech-recognition100K<n<1M1 likes145 downloads5mo agoHugging Face10nyuuzyou /OpenGameArt-GPL-3.0 Dataset Card for OpenGameArt-GPL-3.0 Dataset Summary This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the GNU General Public License version 3.0 (GPL-3.0). The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, and textures along with their associated metadata. Languages The dataset is primarily monolingual: English (en): All asset descriptions… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-GPL-3.0.audioimage-classificationn<1K0 likes138 downloads1y agoHugging Face11irfankabir02 /OpenGameArt-GPL-3.0 Dataset Card for OpenGameArt-GPL-3.0 Dataset Summary This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the GNU General Public License version 3.0 (GPL-3.0). The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, and textures along with their associated metadata. Languages The dataset is primarily monolingual: English (en): All asset descriptions… See the full description on the dataset page: https://huggingface.co/datasets/irfankabir02/OpenGameArt-GPL-3.0.audioimage-classificationn<1K0 likes130 downloads10mo agoHugging Face12Codyfederer /test321 test321 This is a merged speech dataset containing 118 audio segments from 2 source datasets. Dataset Information Total Segments: 118 Speakers: 4 Languages: tr Emotions: happy, angry, sad, neutral Original Datasets: 2 Dataset Structure Each example contains: audio: Audio file (WAV format, 16kHz sampling rate) text: Transcription of the audio speaker_id: Unique speaker identifier (made unique across all merged datasets) emotion: Detected emotion… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/test321.audioautomatic-speech-recognitionn<1K0 likes31 downloads1y agoHugging Face13Rakancorle1 /vggsync-3k VGGSync-3K · out-of-domain audio-visual sync benchmark Out-of-domain evaluation set used in the paper When Vision Speaks for Sound. Derived from VGGSoundSync, this 3,000-clip slice tests whether a video-capable MLLM can detect audio temporal offsets on everyday sound events outside the THUD in-domain training distribution. Each item is one VGGSound clip in one of three conditions: Condition Count gt_synced gt_direction gt_offset_sec Audio aligned (no shift) 1,000 true… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/vggsync-3k.audioaudio-classification1K<n<10K0 likes28 downloads4mo agoHugging Face14malaiwah /qwen3-tts-preset-voices Qwen3-TTS preset voice embeddings The 9 named speakers from Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice packaged as a small sidecar bundle usable with the -Base checkpoint. bundle.safetensors — 9 × 2048-d bfloat16 rows, ~37 KB total bundle.json — metadata (speaker name → spk_id, gender, supported languages) Each row is lifted from talker.model.codec_embedding.weight in the CustomVoice checkpoint at the speaker-ID index from its config.json. With these rows, you can: Deploy only… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qwen3-tts-preset-voices.textn<1K0 likes19 downloads6mo agoHugging Face15instinct-org /yt3_chunked_tokenizedgated yt3_chunked_48k_tokenized This is a gated Russian tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt3_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes17 downloads4mo agoHugging Face16MohamedHussienOmar /whisper-finetune-audio_test3audion<1K0 likes5 downloads1y agoHugging Face17seleznevchamp1 /chester_bennington_30tracks_datasetaudion<1K0 likes5 downloads8mo agoHugging Face18ThisUsernameAlreadyExistsAlreadyExists /aitf-dfk3-synthetic-audio-datasetgatedaudioaudio-classification1K<n<10K0 likes3 downloads6mo agoHugging Face19Mcycm /lntest3audion<1K0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.