datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
triplets_audio_image_text_v1Trio-Image-Audio-Text
Trio
A unified multimodal dataset combining image, audio, and text from diverse public sources.
Usage
This dataset uses Configurations (Subsets) to manage its diverse data sources. You can load specific parts or the entire "filtered" dataset without downloading the NSFW portions.
pip install datasets
1. Load the "filtered" Subset
This configuration loads all 29 safe subsets, excluding the NSFW content.
from datasets import load_dataset
#… See the full description on the dataset page: https://huggingface.co/datasets/VINAY-UMRETHE/Trio-Image-Audio-Text.Emergence-Text-Image-Audio-3D
Emergence: The Four Forms of Intelligence
Summary
A multimodal dataset that unifies Text, Image, Audio, and 3D modalities with quad-modality alignment for every sample, ensuring that each record contains semantically consistent representations of the same concept.
This dataset is curated by using 3D assets from Objaverse as anchors and aligning them with semantically corresponding images and audio clips from various sources using a embedding search… See the full description on the dataset page: https://huggingface.co/datasets/VINAY-UMRETHE/Emergence-Text-Image-Audio-3D.exaple_audio_image_text_triplet
