CoolFace
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Multilingual-Multimodal-NLP /McEvalMcEval benchmark data as described in the McEval Paper. Code for the evaluation can be found on Github as McEval. texttext-generation10K<n<100K21 likes3.9k downloads2y agoHugging Face02ken-sungmin /propagator-multimodal-pretraining-data Propagator Multimodal Pretraining Data This public dataset contains tokenized multimodal pretraining data prepared for the Propagator model family. It combines language, image-grounded, and speech/audio-token examples into a single training format. This is not a raw text or image browsing dataset. The examples have already been converted into compact binary token frames for model training, with a manifest that records the source groups and file layout. Source Code… See the full description on the dataset page: https://huggingface.co/datasets/ken-sungmin/propagator-multimodal-pretraining-data.texttext-generation0 likes1.4k downloads3mo agoHugging Face03cjerzak /MultimodalMathBenchmarks MultimodalMathBenchmarks This repository contains the datasets for the paper Multiplication in Multimodal LLMs: Computation with Text, Image, and Audio Inputs (ACL Findings 2026). It covers the public benchmark datasets and their modality assets (text, images, and audio) used to evaluate the arithmetic capabilities of multimodal LLMs. Canonical Upload Manifest HF path Local source Count Purpose SharedMultimodalGrid.csv SavedData/SharedMultimodalGrid.csv… See the full description on the dataset page: https://huggingface.co/datasets/cjerzak/MultimodalMathBenchmarks.audioimage-text-to-text10K<n<100K0 likes663 downloads2mo agoHugging Face04aliencaocao /multimodal_meme_classification_singapore Dataset Card for Offensive Memes in Singapore Context Dataset Details Dataset Description This dataset is a collection of memes from various existing datasets, online forums, and freshly scrapped contents. It contains both global-context memes and Singapore-context memes, in different splits. It has textual description and a label stating if it is offensive under Singapore society's standards. Curated by: Cao Yuxuan, Wu Jiayang, Alistair Cheong, Theodore Lee… See the full description on the dataset page: https://huggingface.co/datasets/aliencaocao/multimodal_meme_classification_singapore.imagetext-generation100K<n<1M1 likes637 downloads2y agoHugging Face05Baekpica /Inkling-Small-Multimodal-Calibration Inkling-Small Multimodal Calibration The exact 1,663 samples used for BF16 routed-expert importance collection for Inkling-Small Mixed Quant GGUF. This is calibration material, not a held-out evaluation benchmark. The primary balanced pass is: Category Samples Valid decoder tokens Share Text / reasoning 462 471,858 44.976% Code / tool-oriented source text 205 209,715 19.989% Real image / document 486 262,476 25.018% Real speech audio 309 105,080 10.016% Total… See the full description on the dataset page: https://huggingface.co/datasets/Baekpica/Inkling-Small-Multimodal-Calibration.tabulartext-generation1K<n<10K0 likes411 downloads15d agoHugging Face06Multilingual-Multimodal-NLP /McEval-InstructMcEval-Instruct data as described in the McEval Paper. Code for the evaluation and sft can be found on Github as McEval. texttext-generation10K<n<100K39 likes288 downloads2y agoHugging Face07alibaba-multimodal-industrial-ai /IndustryBench IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs 💻Github | 📝Paper IndustryBench is a multi-lingual benchmark for evaluating the industrial domain knowledge of large language models. It comprises 2,049 expert-curated QA pairs spanning 12 industrial sectors, with human-reviewed translations in Chinese, English, Russian, and Vietnamese. Overview Dimension Details Total questions 2,049 Languages Chinese (zh), English (en), Russian (ru)… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench.textquestion-answering1K<n<10K30 likes212 downloads4mo agoHugging Face08superviselab /multimodal-video-annotation-samples Video Annotation Samples – SuperviseLab SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories. Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.tabularvideo-classificationn<1K1 likes119 downloads6mo agoHugging Face09beatsprom /multimodal-vision-language-video-models-2026 👁️ Multimodal Vision-Language & Video Foundation Models Dataset (2026 Edition) A structured research dataset featuring 1,000 domain-verified research papers and code repositories focused on Multimodal Vision-Language Models (VLM), Video Foundation Models, Diffusion Transformers (DiT), Visual Grounding, and World Simulators. Built with Universal Scientific Engine V15.1 Gold, providing 47 schema attributes with verified repository attribution, modality capability matrix, vision… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-language-video-models-2026.tabularfeature-extractionn<1K3 likes107 downloads1mo agoHugging Face10beatsprom /multimodal-vision-ocr-document-parsing-2026 📐 Multimodal Vision-Language & Industrial OCR Document Parsing SFT/DPO Suite (2026) This repository provides the official 100-sample production teaser of the Multimodal Vision-Language & Industrial OCR Document Parsing SFT/DPO Suite (2026) by BeatsProm AI Research Lab. The dataset is engineered to train open-weights Vision-Language Models (Qwen2-VL, Pixtral-12B, Llama-3.2-Vision, ColPali) on dense document parsing, normalized spatial bounding boxes (<box>[ymin, xmin, ymax… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-ocr-document-parsing-2026.textimage-to-textn<1K0 likes88 downloads21d agoHugging Face11neurips26 /MultimodalUnlearningEvalBenchmark 🧠 Multimodal Unlearning Evaluation Benchmark 📌 Overview This dataset provides evaluation outputs for studying metric inconsistency in multimodal machine unlearning. It supports reproducibility of results in: Metric Unreliability in Multimodal Machine Unlearning (NeurIPS 2026) 📊 Contents File Description 📄 multimodal_results.json Results on VQA benchmarks (MLLMU-Bench, UnLOK-VQA, MMUBench) 📄 unimodal_results.json CIFAR-10… See the full description on the dataset page: https://huggingface.co/datasets/neurips26/MultimodalUnlearningEvalBenchmark.tabulartext-generationn<1K1 likes82 downloads5mo agoHugging Face12lv12 /MultiModalDataset Dataset Card for MultiModal Dataset Dataset Description Dataset Summary MultiModal Dataset is a curated collection of 85,000 samples spanning three modalities: text, images, and audio. It combines high-quality web content, image-caption pairs from COCO 2017, and audio samples from AudioSet to enable comprehensive multimodal model training and evaluation. The dataset is organized into three subsets: fineweb: 37,500 high-quality web text samples (>8… See the full description on the dataset page: https://huggingface.co/datasets/lv12/MultiModalDataset.imagetext-generation10K<n<100K0 likes80 downloads1mo agoHugging Face13matlok /multimodal-python-copilot-training-overview Multimodal Datasets for Training Python Copilots from Source Code Analysis Welcome to the matlok multimodal python copilot training datasets. This is an overview for our training and fine-tuning datasets found below: ~2.3M unique source coding rows 1.1M+ instruct alpaca yaml text rows updated bi-weekly ~923K png knowledge graph images with alpaca text description ~334K mp3s over ~2 years of continuous audio playtime requires 1.5 TB storage on disk Please reach out if you find an… See the full description on the dataset page: https://huggingface.co/datasets/matlok/multimodal-python-copilot-training-overview.texttext-generation10K<n<100K28 likes63 downloads3y agoHugging Face14syhuggingface /multimodal_rewardbench Dataset Card for Multimodal RewardBench 🏆 Dataset Attribution This dataset is created by Yasunaga et al. (2025). 📄 Paper: Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models 💻 GitHub Repository: https://github.com/facebookresearch/multimodal_rewardbench I have downloaded the dataset from the GitHub repo and only modified the "Image" attribute by converting file paths to datasets.Image() for easier integration with 🤗… See the full description on the dataset page: https://huggingface.co/datasets/syhuggingface/multimodal_rewardbench.imageimage-to-text1K<n<10K0 likes57 downloads2y agoHugging Face15alst10 /gemma4-multimodal-recipe-dataset 🍳 Gemma 4 Multimodal Recipe & Food Dataset A balanced, high-density multimodal dataset curated specifically for fine-tuning compact vision-language models (such as gemma-4-e2b-it) for visual food recognition, recipe generation, and dietary recommendation. 🔗 Upstream & Source Datasets This dataset was created by cleaning, reformatting, and synthesizing samples across the following 5 Hugging Face sources: Dataset Modality Role in Pipeline… See the full description on the dataset page: https://huggingface.co/datasets/alst10/gemma4-multimodal-recipe-dataset.imagevisual-question-answering10K<n<100K0 likes52 downloads1mo agoHugging Face16emgena /multimodal_rag_complex_table_extractor_teaser 🚀 Data Platform - Multi-Modal RAG, Complex Document & Table Extractor (Evaluation Teaser) ⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (500 Samples) & Commercial EULA on Gumroad:👉 Data Platform - Multi-Modal RAG, Complex Document & Table Extractor on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout! 📦 What is Inside the Full Production Package: 500 Verified FAANG v2.0 Scenarios (100%… See the full description on the dataset page: https://huggingface.co/datasets/emgena/multimodal_rag_complex_table_extractor_teaser.texttext-generationn<1K0 likes41 downloads6d agoHugging Face17ai2lumos /lumos_multimodal_ground_iterative 🪄 Agent Lumos: Unified and Modular Training for Open-Source Language Agents 🌐[Website]   📝[Paper]   🤗[Data]   🤗[Model]   🤗[Demo]   We introduce 🪄Lumos, Language Agents with Unified Formats, Modular Design, and Open-Source LLMs. Lumos unifies a suite of complex interactive tasks and achieves competitive performance with GPT-4/3.5-based and larger open-source agents. Lumos has following features: 🧩 Modular Architecture: 🧩 Lumos consists of planning, grounding… See the full description on the dataset page: https://huggingface.co/datasets/ai2lumos/lumos_multimodal_ground_iterative.texttext-generation10K<n<100K2 likes40 downloads3y agoHugging Face18AYI-NEDJIMI /ai-code-multimodal-fr Dataset : IA Generation de Code, IA Multimodale & Small Language Models (FR) Description Dataset francophone couvrant les outils d'assistance au codage par IA, l'IA multimodale, les Small Language Models (SLM) et GraphRAG. Ce dataset est concu pour la recherche, la formation et le developpement d'applications dans le domaine de l'IA appliquee au developpement logiciel et a la cybersecurite. Articles couverts IA pour la Generation de Code : Copilot, Cursor… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ai-code-multimodal-fr.textquestion-answeringn<1K0 likes40 downloads7mo agoHugging Face19Truthseeker87 /solarhive-community-solar-multimodal SolarHive Community Solar Dataset Canonical training corpus for the SolarHive family of fine-tuned Gemma 4 models. 1,727 rows (1,713 text + 14 image-grounded). A combined text + sky-image training corpus for community solar energy intelligence. Built to fine-tune Gemma 4 into an AI energy advisor for residential solar microgrids — answering questions about production, storage, grid mix, weather impact, maintenance scheduling, and cross-source planning, with native… See the full description on the dataset page: https://huggingface.co/datasets/Truthseeker87/solarhive-community-solar-multimodal.imagequestion-answering1K<n<10K0 likes40 downloads5mo agoHugging Face20AYI-NEDJIMI /ai-code-multimodal-en Dataset: AI Code Generation, Multimodal AI & Small Language Models (EN) Description English dataset covering AI coding assistants, multimodal AI, Small Language Models (SLMs), and GraphRAG. This dataset is designed for research, training, and application development in the field of AI applied to software development and cybersecurity. Articles Covered AI Code Generation: Copilot, Cursor, Claude Code - Comparison of 12 leading AI coding assistants Computer… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ai-code-multimodal-en.textquestion-answeringn<1K0 likes39 downloads7mo agoHugging Face21ai2lumos /lumos_multimodal_plan_iterative 🪄 Agent Lumos: Unified and Modular Training for Open-Source Language Agents 🌐[Website]   📝[Paper]   🤗[Data]   🤗[Model]   🤗[Demo]   We introduce 🪄Lumos, Language Agents with Unified Formats, Modular Design, and Open-Source LLMs. Lumos unifies a suite of complex interactive tasks and achieves competitive performance with GPT-4/3.5-based and larger open-source agents. Lumos has following features: 🧩 Modular Architecture: 🧩 Lumos consists of planning, grounding… See the full description on the dataset page: https://huggingface.co/datasets/ai2lumos/lumos_multimodal_plan_iterative.texttext-generation10K<n<100K2 likes38 downloads3y agoHugging Face22obaydata /svg-multimodal-rubrics SVG Multimodal Rubrics A multimodal dataset of SVG code generation samples with natural language descriptions and evaluation rubrics. Each sample pairs a detailed prompt (Markdown) with its corresponding SVG source code, covering animations, 3D scenes, games, and visual effects. Designed for training and evaluating models on visual code generation — generating complex, interactive SVG artwork from natural language descriptions. Overview Item Details Samples… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/svg-multimodal-rubrics.imagetext-generationn<1K0 likes35 downloads6mo agoHugging Face23lihicarmeli /fashion-stylist-multimodal 👗 Fashion Stylist Multimodal Dataset A synthetic multimodal dataset pairing structured fashion metadata, styled outfit text, and AI-generated portraits. 🎯 Overview This dataset contains ~1,000 synthetic fashion-styling profiles, each combining: 🧬 Structured demographic & style metadata 📝 A styled outfit description with e-commerce search queries 🖼️ A generated 512×512 studio-style portrait of a fictional person wearing the outfit The dataset was built… See the full description on the dataset page: https://huggingface.co/datasets/lihicarmeli/fashion-stylist-multimodal.imagetext-to-image1K<n<10K0 likes31 downloads1mo agoHugging Face24multimodal-reframing /mirror MIRROR Dataset MIRROR is a synthetic vision–language dataset for multimodal cognitive reframing under client resistance. Paper: 🪞 MIRROR: Multimodal Cognitive Reframing Therapy for Rolling with Resistance The dataset includes: Client profile metadata (CACTUS idx, CelebA idx) Dialogue written in a screenplay format, including stage directions that describe facial expressions ⚠️ Images themselves are not included to comply with the CelebA license. However, we provide the full image… See the full description on the dataset page: https://huggingface.co/datasets/multimodal-reframing/mirror.tabulartext-generationn<1K2 likes24 downloads10mo agoHugging Face25Yzineb /arabic_multimodal_text Arabic Multimodal Text Dataset This dataset contains Arabic text extracted from public domain books available on Archive.org. Dataset Details Language: Arabic Total Text Samples: 162 Source PDFs: 11 out of 50 Format: JSONL Structure Each line in the JSONL file contains: { "source": "URL of the original PDF", "title": "Title of the book", "text": "Extracted Arabic text" } Usage from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Yzineb/arabic_multimodal_text.texttext-generationn<1K1 likes22 downloads11mo agoHugging Face26Orib24 /Roomly-Student-Bios-Multimodal Roomly: Multimodal Roommate Matching Dataset 🎯 Problem Statement Finding a roommate is often reduced to dry filters like "budget" and "location". Roomly aims to revolutionize this by focusing on personality, lifestyle, and visual preferences. This dataset provides synthetic student profiles and their ideal room environments. 📊 Exploratory Data Analysis (EDA) 1. User Persona Distribution Our dataset contains a balanced mix of different student… See the full description on the dataset page: https://huggingface.co/datasets/Orib24/Roomly-Student-Bios-Multimodal.texttext-generationn<1K0 likes6 downloads9mo agoHugging Face27audibeal74 /panta_instruct_multi_modal_v1 Panta Instruct Multi-Modal v1 Dataset d'instructions multimodal en français : chaque exemple associe une question (texte + parole + pictogrammes) à une réponse (texte + pictogrammes). Colonnes Colonne Type Description audio Audio (24 kHz, mono) Enregistrement de la question (text_input) text_input string Question / instruction text_output string Réponse pictos_input list[string] Identifiants des pictogrammes de la question pictos_output… See the full description on the dataset page: https://huggingface.co/datasets/audibeal74/panta_instruct_multi_modal_v1.audioautomatic-speech-recognition10K<n<100K0 likes16h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.