datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sarab
Sarab Dataset
The dataset behind Sarab, a cause-diagnostic Arabic visual hallucination
evaluation benchmark for multimodal LLMs, modeled on Liu et al.'s CVPR 2025 PhD
benchmark. Code and evaluation scripts are on
GitHub.
What this is
A human-captioned pool of Arabic Cultural Visual Vocabulary (ACVV) images
(architecture, attire, cuisine, objects, script), built into five evaluation
modes:
base — plain image, direct Arabic question.
sec (specious context) — image… See the full description on the dataset page: https://huggingface.co/datasets/Sarab-MLLMs/sarab.mllm-self-fullfillingMLLM-SEG-resultsMLLMstage3-textcaps-scienceqa-textvqaMLLMsMLLM_SAE_TRAIN
