CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01flax-community /conceptual-12m-mbart-50-multilingualimage10M<n<100M2 likes521 downloads5y agoHugging Face02Multimedia-SMU /seeingculture-benchmarkPaper | Project Page | Leaderboard | Explorer | Code | CMB, the video successor Seeing Culture Benchmark (SCB) Evaluating Visual Reasoning and Grounding in Cultural Context The Seeing Culture Benchmark (SCB) evaluates cultural reasoning in vision-language models in two stages: i) selecting the correct visual option with multiple-choice visual question answering (VQA), and ii) segmenting the relevant cultural artifact as evidence of reasoning. Visual options in… See the full description on the dataset page: https://huggingface.co/datasets/Multimedia-SMU/seeingculture-benchmark.imageimage-text-to-text1K<n<10K3 likes138 downloads16h agoHugging Face03flax-community /conceptual-12m-multilingual-marian-esimage1M<n<10M0 likes84 downloads3y agoHugging Face04illilliks /multimodal-time-series-forecastingimage100K<n<1M2 likes49 downloads1y agoHugging Face05hsiangfu /multimodal_query_rewrites ReVision: Visual Instruction Rewriting Dataset Dataset Summary The ReVision dataset is a large-scale collection of task-oriented multimodal instructions, designed to enable on-device, privacy-preserving Visual Instruction Rewriting (VIR). The dataset consists of 39,000+ examples across 14 intent domains, where each example comprises: Image: A visual scene containing relevant information. Original instruction: A multimodal command (e.g., a spoken query referencing visual… See the full description on the dataset page: https://huggingface.co/datasets/hsiangfu/multimodal_query_rewrites.image10K<n<100K0 likes40 downloads2y agoHugging Face06anonymoususerrevision /multimodal_query_rewrites ReVision: Visual Instruction Rewriting Dataset Dataset Summary The ReVision dataset is a large-scale collection of task-oriented multimodal instructions, designed to enable on-device, privacy-preserving Visual Instruction Rewriting (VIR). The dataset consists of 39,000+ examples across 14 intent domains, where each example comprises: Image: A visual scene containing relevant information. Original instruction: A multimodal command (e.g., a spoken query referencing visual… See the full description on the dataset page: https://huggingface.co/datasets/anonymoususerrevision/multimodal_query_rewrites.image10K<n<100K1 likes34 downloads2y agoHugging Face07JourneyBench /JourneyBench_Multi_Image_VQA Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/JourneyBench/JourneyBench_Multi_Image_VQA.imagequestion-answeringn<1K0 likes23 downloads2y agoHugging Face08ruslanmv /hotel-multimodalimage1M<n<10M2 likes8 downloads2y agoHugging Face09GRAI-UNSTPB /MultiFakeRomimagen<1K0 likes6 downloads4mo agoHugging Face10Shoriful025 /image_caption_pairs_for_multimodalimagen<1K0 likes1 downloads9mo agoHugging Face11KotaroOmote /rg-7wildlife-multilabel-v1image1K<n<10K0 likes1 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.