CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kkkyao /triplets_audio_image_text_v1audio1K<n<10K0 likes57 downloads9mo agoHugging Face02VINAY-UMRETHE /Trio-Image-Audio-Textgated Trio A unified multimodal dataset combining image, audio, and text from diverse public sources. Usage This dataset uses Configurations (Subsets) to manage its diverse data sources. You can load specific parts or the entire "filtered" dataset without downloading the NSFW portions. pip install datasets 1. Load the "filtered" Subset This configuration loads all 29 safe subsets, excluding the NSFW content. from datasets import load_dataset #… See the full description on the dataset page: https://huggingface.co/datasets/VINAY-UMRETHE/Trio-Image-Audio-Text.audio100K<n<1M0 likes29 downloads7mo agoHugging Face03VINAY-UMRETHE /Emergence-Text-Image-Audio-3Dgated Emergence: The Four Forms of Intelligence Summary A multimodal dataset that unifies Text, Image, Audio, and 3D modalities with quad-modality alignment for every sample, ensuring that each record contains semantically consistent representations of the same concept. This dataset is curated by using 3D assets from Objaverse as anchors and aligning them with semantically corresponding images and audio clips from various sources using a embedding search… See the full description on the dataset page: https://huggingface.co/datasets/VINAY-UMRETHE/Emergence-Text-Image-Audio-3D.textany-to-any10K<n<100K0 likes20 downloads7mo agoHugging Face04younissk /audio-text-embed-to-imagesimagen<1K0 likes11 downloads1y agoHugging Face05LeroyDyer /Spectrogram_Audio_text_to_Base64imagen<1K2 likes10 downloads2y agoHugging Face06LeroyDyer /Bass_Audio_text_to_Base64imagen<1K1 likes6 downloads2y agoHugging Face07davanstrien /exaple_audio_image_text_tripletaudion<1K0 likes5 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.