MLLMs
Datasets
All datasets matching “MLLMs”sarab
Sarab Dataset
The dataset behind Sarab, a cause-diagnostic Arabic visual hallucination
evaluation benchmark for multimodal LLMs, modeled on Liu et al.'s CVPR 2025 PhD
benchmark. Code and evaluation scripts are on
GitHub.
What this is
A human-captioned pool of Arabic Cultural Visual Vocabulary (ACVV) images
(architecture, attire, cuisine, objects, script), built into five evaluation
modes:
base — plain image, direct Arabic question.
sec (specious context) — image… See the full description on the dataset page: https://huggingface.co/datasets/Sarab-MLLMs/sarab.mllm-shap
MLLM-SHAP experiment datasets
Curated test splits for studying Shapley-value explanations in multimodal large language models (text and audio inputs). Each configuration is a filtered, size-controlled subset built for reproducible benchmarking—not a full copy of the upstream corpora.
Configs follow the naming pattern {task}__{source} (for example single_sentence__voice_bench).
Quick load
Pin a dataset revision for reproducibility (replace REVISION with the commit hash… See the full description on the dataset page: https://huggingface.co/datasets/Pawlo77/mllm-shap.Hallucination_MLLMs_Data_preprocessingmllm-shap-new
MLLM Shap Experiments Datasets
This repo contains following datasets used in modality and multilinguality experiments using Shapley Values, including:
Datasets based on Infinity Instruct Dataset:
Multi Lingual Dataset - 99 rows (33 in english, 33 in french, 33 in spanish), all of them translated to remaining 2 languages - resulting in total of 297 rows.
Datasets based on Voice Bench Dataset
Multi Sentence Dataset - 250 rows in English, multi-sentence entries in english.
Single… See the full description on the dataset page: https://huggingface.co/datasets/mvishiu11/mllm-shap-new.mllm-self-fullfillingmllm-shap-copy
MLLM Shap Experiments Datasets
This repo contains following datasets used in modality and multilinguality experiments using Shapley Values, including:
Datasets based on Infinity Instruct Dataset:
Multi Lingual Dataset - 99 rows (33 in english, 33 in french, 33 in spanish), all of them translated to remaining 2 languages - resulting in total of 297 rows.
Datasets based on Voice Bench Dataset
Multi Sentence Dataset - 250 rows in English, multi-sentence entries in english.
Single… See the full description on the dataset page: https://huggingface.co/datasets/mvishiu11/mllm-shap-copy.
