reality-check
reality-check-on-context-utilisation
Dataset card for the dataset used in "A Reality Check on Context Utilisation for Retrieval-Augmented Generation"
Dataset Details
This dataset was used for the analysis and plots in the paper "A Reality Check on Context Utilisation for Retrieval-Augmented Generation". More details on the dataset can be found in the paper.
Dataset Description
The dataset contains samples from CounterFact (Meng et al. 2022), ConflictQA (Xie et al. 2024), and DRUID with… See the full description on the dataset page: https://huggingface.co/datasets/copenlu/reality-check-on-context-utilisation.LLMEval-Reality_CheckThis repository contains model evaluation results with the following setup:
Models evaluated:
GPT-5
Gemini-2.5-Pro
Claude-4.5-Sonnet
Datasets included:
MMLU-Pro (test split)
GPQA (three subsets)
MATH-500
MMMU-Pro (standard 10 options and vision versions)
Each split in the dataset corresponds to one benchmark.
Schema
All datasets have been standardized to a unified schema with the following features:
dataset_info:
features:
- id: int64
- prompt: string… See the full description on the dataset page: https://huggingface.co/datasets/HappyEval/LLMEval-Reality_Check.vlm-reality-check
VLM Reality Check: A Causal-Contrastive Benchmark for Vision-Language Models
VLM Reality Check is a large-scale, automatically generated dataset designed to evaluate the causal reasoning and bias resistance of Vision-Language Models (VLMs). It features nearly 100,000 challenges across 14 bias types, utilizing counterfactual image transforms to create minimal-edit contrastive pairs.
Dataset Summary
Total Challenges: 95,317
Source Images: 1,868 (from Kaggle/ECCV v3… See the full description on the dataset page: https://huggingface.co/datasets/anonymous9457/vlm-reality-check.
