CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Ardea /NEXUS-temporal_hierarchical_multi-modal NEXUS: Neural Evolution for eXtensible Universal Semantics Dataset (Temporal Multimodal Slices) This dataset is a multi-modal, hierarchical, temporal representation derived from HuggingFaceFV/finevideo. It is designed for streaming training where the primary unit is a 10 ms "slice" that aggregates upward into moments (100 ms), seconds (1 s), experiences (10 s), and minutes (60 s). It is meant to represent an extensible stream of "experience" as there are… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/NEXUS-temporal_hierarchical_multi-modal.imageautomatic-speech-recognition10M<n<100M5 likes888 downloads4mo agoHugging Face02zwq2018 /Multi-modal-Self-instruct Dataset Description Paper Information Dataset Examples Leaderboard Dataset Usage Data Downloading Data Format Evaluation Citation You can download the zip dataset directly, and both train and test subsets are collected in Multi-modal-Self-instruct.zip. Dataset Description Multi-Modal Self-Instruct dataset utilizes large language models and their code capabilities to synthesize massive abstract images and visual reasoning instructions across daily scenarios. This benchmark… See the full description on the dataset page: https://huggingface.co/datasets/zwq2018/Multi-modal-Self-instruct.imagemultiple-choice10K<n<100K34 likes558 downloads2y agoHugging Face03mixed-modality-search /MixBench2026 MixBench: A Benchmark for Mixed Modality Retrieval MixBench is a benchmark for evaluating retrieval across text, images, and multimodal documents. It is designed to test how well retrieval models handle queries and documents that span different modalities, such as pure text, pure images, and combined image+text inputs. MixBench includes four subsets, each curated from a different data source: MSCOCO Google_WIT VisualNews OVEN Each subset contains: queries.jsonl: each entry… See the full description on the dataset page: https://huggingface.co/datasets/mixed-modality-search/MixBench2026.imagetext-ranking10K<n<100K1 likes276 downloads8mo agoHugging Face04mteb /coco-modality-equivalenceaudio1K<n<10K0 likes142 downloads13d agoHugging Face05rakshi719 /coco-modality-equivalenceaudio1K<n<10K0 likes101 downloads20d agoHugging Face06ModalityDance /Optical-Reasoning-4k Optical Reasoning Overview Optical Reasoning contains 3,907 rendered visual rationales used in "Optical Reasoning: Rethinking Images as an Expressive Reasoning Medium Beyond Text". It covers 5 benchmarks, including typographic rationales for all benchmarks and graphical rationales for AQuA-RAT. AQuA-RAT: Multiple-choice algebra and quantitative reasoning problems with five answer options. GPQA Diamond: Graduate-level multiple-choice science questions spanning… See the full description on the dataset page: https://huggingface.co/datasets/ModalityDance/Optical-Reasoning-4k.imagevisual-question-answering1K<n<10K0 likes58 downloads4mo agoHugging Face07vlm-modality-research /gsm8k-rendered-vlm-v2 GSM8K Rendered-VL v2 1319 rendered GSM8K test problems for the VLM modality study (Phase 1). Contributors Rodela Ghosh — study design, pilot (v1), dataset packaging and Hugging Face release (scripts/prepare_hf_v2_release.py) Aviral Gupta — benchmark infrastructure (src/), v2 rendering protocol (src/rendering.py), Phase 1 model runs Code: https://github.com/Ro-netizen004/vlm-modality-research Not interchangeable with v1: RodelaG/gsm8k-rendered-vlm v1 v2… See the full description on the dataset page: https://huggingface.co/datasets/vlm-modality-research/gsm8k-rendered-vlm-v2.image1K<n<10K0 likes57 downloads3mo agoHugging Face08ModalityDance /AR-Omni-Instruct-v0.1 AR-Omni-Instruct Overview AR-Omni-Instruct is a multimodal instruction-tuning dataset for training unified autoregressive any-to-any models. All modalities are represented as discrete tokens in a single interleaved token stream, enabling standard next-token prediction training over multimodal sequences. Dataset Summary Type: multimodal instruction-tuning data Format: discrete tokenized multimodal conversations / sequences Use case: instruction tuning… See the full description on the dataset page: https://huggingface.co/datasets/ModalityDance/AR-Omni-Instruct-v0.1.audioany-to-any100K<n<1M0 likes52 downloads8mo agoHugging Face09anya-ji /multi-modal-image-editimage1K<n<10K1 likes44 downloads1y agoHugging Face10vlm-modality-research /modality-conflict-arbitration-v2 Modality-Conflict Arbitration Benchmark (v2) A controlled benchmark for studying how a vision-language model arbitrates between its two input channels when they disagree — and whether that choice tracks the reliability of each channel. Each row is a single conflict trial: an image of one math problem paired with the text of a different problem. Because the two ground-truth answers are carried side by side, the model's output alone tells you which modality it followed — no… See the full description on the dataset page: https://huggingface.co/datasets/vlm-modality-research/modality-conflict-arbitration-v2.imagevisual-question-answering10K<n<100K0 likes42 downloads2mo agoHugging Face11bombshelll /brain_modalityimage1K<n<10K1 likes35 downloads2y agoHugging Face12ESmike /true_imgs_modalities Dataset Summary ESmike/true_imgs_modalities is a multimodal dataset containing 4,149 real images at a uniform resolution of 512 × 512, along with multiple derived modalities for each image. This dataset is intended for research and experimentation in computer vision, generative modeling, and multimodal learning. All images originate from real photographs curated and processed. Each base image is provided alongside additional modality representations (details in the Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/ESmike/true_imgs_modalities.imageimage-to-image1K<n<10K2 likes29 downloads10mo agoHugging Face13vlm-modality-research /chartqa-evidence-conflict-v2 ChartQA-Conflict v2 This dataset contains 230 reviewed conflicts between a native ChartQA chart and an evidence-bearing textual report. The original chart supports one answer, while counterfactual facts in the report support a distinct answer to the same question. Neither source is privileged in the evaluation prompt. Each row contains the original chart, shared question, chart-supported answer, report-supported answer, evidence-bearing report, unit and counterfactual strategy… See the full description on the dataset page: https://huggingface.co/datasets/vlm-modality-research/chartqa-evidence-conflict-v2.imagevisual-question-answeringn<1K0 likes28 downloads2mo agoHugging Face14dingyue1011 /Modality-Alignimage1K<n<10K0 likes27 downloads2y agoHugging Face15mesolitica /translate-Multi-modal-Self-instruct Translated https://huggingface.co/datasets/zwq2018/Multi-modal-Self-instruct Translate to Malay using https://mesolitica.com/translation Base model, a nice dataset for visual QA charts, tables, simulated maps, dashboards, flowcharts, relation graphs, floor plans, and visual puzzles. image10K<n<100K2 likes26 downloads2y agoHugging Face16vlm-modality-research /chartqa-evidence-conflict-table-v1 ChartQA-Conflict chart/table representation ablation This derivative contains 229 audited ChartQA-Conflict items. Each row provides two visual renderings of the same official ChartQA facts: chart_image: the original ChartQA chart; table_image: a plain table image rendered from the corresponding official ChartQA CSV. The shared question, chart-supported answer, evidence-bearing conflicting report, report-supported answer, and Source A/B assignment are identical across the two… See the full description on the dataset page: https://huggingface.co/datasets/vlm-modality-research/chartqa-evidence-conflict-table-v1.imagevisual-question-answeringn<1K0 likes26 downloads2mo agoHugging Face17anonymous052025 /multimodal-modality-conflict-datasetimage10K<n<100K0 likes25 downloads1y agoHugging Face18Ibrahim-Alam /multi-modal_offensive_memeimagen<1K0 likes18 downloads8mo agoHugging Face19hyokwan /multi_modal_sampleimagen<1K0 likes15 downloads4mo agoHugging Face20joshbarua /conflicting_modalityimage1K<n<10K0 likes14 downloads6mo agoHugging Face21depgh4311 /heroicons_modalimagen<1K0 likes13 downloads2y agoHugging Face22hyokwan /multi_modal_sportsimagen<1K0 likes11 downloads6mo agoHugging Face23vlm-modality-research /math-rendered-vlm-v1imagen<1K0 likes11 downloads3mo agoHugging Face24vlm-modality-research /chartqa-evidence-conflict-v1 ChartQA Evidence Conflict This dataset contains 230 reviewed conflicts between a native ChartQA chart and an evidence-bearing textual report. The original chart supports one answer, while counterfactual facts in the report support a distinct answer to the same question. Neither source is privileged in the evaluation prompt. Each row contains the original chart, shared question, chart-supported answer, report-supported answer, evidence-bearing report, unit and counterfactual… See the full description on the dataset page: https://huggingface.co/datasets/vlm-modality-research/chartqa-evidence-conflict-v1.imagevisual-question-answeringn<1K0 likes11 downloads2mo agoHugging Face25patrickamadeus /modality-imbalens-circlesimage10K<n<100K0 likes10 downloads1y agoHugging Face26hcw0329 /multi_modal_sampleimagen<1K0 likes10 downloads3mo agoHugging Face27vlm-modality-research /aqua-rat-rendered-vlm-v1imagen<1K0 likes10 downloads3mo agoHugging Face28hyokwan /multi_modalimagen<1K0 likes9 downloads6mo agoHugging Face29awrub93 /modal-test-dataset Dataset Card for "modal-test-dataset" More Information needed imagen<1K0 likes8 downloads2y agoHugging Face30nirmalendu01 /modality-conflict-datasetimage10K<n<100K0 likes7 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.