CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ken-sungmin /propagator-multimodal-pretraining-data Propagator Multimodal Pretraining Data This public dataset contains tokenized multimodal pretraining data prepared for the Propagator model family. It combines language, image-grounded, and speech/audio-token examples into a single training format. This is not a raw text or image browsing dataset. The examples have already been converted into compact binary token frames for model training, with a manifest that records the source groups and file layout. Source Code… See the full description on the dataset page: https://huggingface.co/datasets/ken-sungmin/propagator-multimodal-pretraining-data.texttext-generation0 likes1.3k downloads3mo agoHugging Face02aliencaocao /multimodal_meme_classification_singapore Dataset Card for Offensive Memes in Singapore Context Dataset Details Dataset Description This dataset is a collection of memes from various existing datasets, online forums, and freshly scrapped contents. It contains both global-context memes and Singapore-context memes, in different splits. It has textual description and a label stating if it is offensive under Singapore society's standards. Curated by: Cao Yuxuan, Wu Jiayang, Alistair Cheong, Theodore Lee… See the full description on the dataset page: https://huggingface.co/datasets/aliencaocao/multimodal_meme_classification_singapore.imagetext-generation100K<n<1M1 likes705 downloads2y agoHugging Face03cjerzak /MultimodalMathBenchmarks MultimodalMathBenchmarks This repository contains the datasets for the paper Multiplication in Multimodal LLMs: Computation with Text, Image, and Audio Inputs (ACL Findings 2026). It covers the public benchmark datasets and their modality assets (text, images, and audio) used to evaluate the arithmetic capabilities of multimodal LLMs. Canonical Upload Manifest HF path Local source Count Purpose SharedMultimodalGrid.csv SavedData/SharedMultimodalGrid.csv… See the full description on the dataset page: https://huggingface.co/datasets/cjerzak/MultimodalMathBenchmarks.audioimage-text-to-text10K<n<100K0 likes669 downloads2mo agoHugging Face04lv12 /MultiModalDataset Dataset Card for MultiModal Dataset Dataset Description Dataset Summary MultiModal Dataset is a curated collection of 85,000 samples spanning three modalities: text, images, and audio. It combines high-quality web content, image-caption pairs from COCO 2017, and audio samples from AudioSet to enable comprehensive multimodal model training and evaluation. The dataset is organized into three subsets: fineweb: 37,500 high-quality web text samples (>8… See the full description on the dataset page: https://huggingface.co/datasets/lv12/MultiModalDataset.imagetext-generation10K<n<100K0 likes79 downloads1mo agoHugging Face05syhuggingface /multimodal_rewardbench Dataset Card for Multimodal RewardBench 🏆 Dataset Attribution This dataset is created by Yasunaga et al. (2025). 📄 Paper: Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models 💻 GitHub Repository: https://github.com/facebookresearch/multimodal_rewardbench I have downloaded the dataset from the GitHub repo and only modified the "Image" attribute by converting file paths to datasets.Image() for easier integration with 🤗… See the full description on the dataset page: https://huggingface.co/datasets/syhuggingface/multimodal_rewardbench.imageimage-to-text1K<n<10K0 likes57 downloads2y agoHugging Face06alst10 /gemma4-multimodal-recipe-dataset 🍳 Gemma 4 Multimodal Recipe & Food Dataset A balanced, high-density multimodal dataset curated specifically for fine-tuning compact vision-language models (such as gemma-4-e2b-it) for visual food recognition, recipe generation, and dietary recommendation. 🔗 Upstream & Source Datasets This dataset was created by cleaning, reformatting, and synthesizing samples across the following 5 Hugging Face sources: Dataset Modality Role in Pipeline… See the full description on the dataset page: https://huggingface.co/datasets/alst10/gemma4-multimodal-recipe-dataset.imagevisual-question-answering10K<n<100K0 likes52 downloads1mo agoHugging Face07lihicarmeli /fashion-stylist-multimodal 👗 Fashion Stylist Multimodal Dataset A synthetic multimodal dataset pairing structured fashion metadata, styled outfit text, and AI-generated portraits. 🎯 Overview This dataset contains ~1,000 synthetic fashion-styling profiles, each combining: 🧬 Structured demographic & style metadata 📝 A styled outfit description with e-commerce search queries 🖼️ A generated 512×512 studio-style portrait of a fictional person wearing the outfit The dataset was built… See the full description on the dataset page: https://huggingface.co/datasets/lihicarmeli/fashion-stylist-multimodal.imagetext-to-image1K<n<10K0 likes47 downloads1mo agoHugging Face08obaydata /svg-multimodal-rubrics SVG Multimodal Rubrics A multimodal dataset of SVG code generation samples with natural language descriptions and evaluation rubrics. Each sample pairs a detailed prompt (Markdown) with its corresponding SVG source code, covering animations, 3D scenes, games, and visual effects. Designed for training and evaluating models on visual code generation — generating complex, interactive SVG artwork from natural language descriptions. Overview Item Details Samples… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/svg-multimodal-rubrics.imagetext-generationn<1K0 likes44 downloads6mo agoHugging Face09Truthseeker87 /solarhive-community-solar-multimodal SolarHive Community Solar Dataset Canonical training corpus for the SolarHive family of fine-tuned Gemma 4 models. 1,727 rows (1,713 text + 14 image-grounded). A combined text + sky-image training corpus for community solar energy intelligence. Built to fine-tune Gemma 4 into an AI energy advisor for residential solar microgrids — answering questions about production, storage, grid mix, weather impact, maintenance scheduling, and cross-source planning, with native… See the full description on the dataset page: https://huggingface.co/datasets/Truthseeker87/solarhive-community-solar-multimodal.imagequestion-answering1K<n<10K0 likes40 downloads5mo agoHugging Face10xaddh /multimodal-privacy Auditing M-LLMs for Privacy Risks: A Synthetic Benchmark and Evaluation Framework Recent advances in multi-modal Large Language Models (M-LLMs) have demonstrated a powerful ability to synthesize implicit information from disparate sources, including images and text. These resourceful data from social media also introduce a significant and underexplored privacy risk: the inference of sensitive personal attributes from seemingly daily media content. However, the lack of benchmarks and… See the full description on the dataset page: https://huggingface.co/datasets/xaddh/multimodal-privacy.imagequestion-answering1K<n<10K1 likes29 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.