datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ChinaHeritaQA
Images
This folder contains visual data for the ChinaHeritaQA benchmark: https://arxiv.org/abs/2606.08959
Contents
Folder
Description
Image_data/
Chinese UNESCO World Heritage Site images (2,279 images from 51 sites)
worlds_data/
Non-Chinese World Heritage Site images (133 images from 23 sites)
Overview
The image dataset includes a comprehensive collection of photographs from both Chinese and international UNESCO World Heritage… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-NLP/ChinaHeritaQA.CAVE
Dataset Card for CAVE: Commonsense Anomalies in Visual Environments
🏠 Project Page📄 Paper (EMNLP 2025)💻 Code
Dataset Details
Dataset Description
CAVE is the first benchmark of real-world visual anomalies for evaluating Vision-Language Models (VLMs). It is curated from images captured in real-life settings (photographs and screenshots taken by individuals), sourced from Reddit.
The benchmark is grounded in cognitive science literature on how humans detect and… See the full description on the dataset page: https://huggingface.co/datasets/epfl-nlp/CAVE.tamawalt-n-imZZyann
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Tamazight-NLP/tamawalt-n-imZZyann.
