AIBE
Datasets
All datasets matching “AIBE”M-VQA
M-VQA: Visual Question Answering under Image Distortions
Dataset Description
M-VQA is the official dataset of the M-VQA Challenge, an ACM Multimedia 2026 Grand Challenge on evaluating visual question answering systems under image distortions.
The dataset contains 12,400 image-question-answer samples:
400 original images;
12,000 distorted images generated from the original images using 30 distortion types;
five distortion severity levels, numbered from 4 to 8.… See the full description on the dataset page: https://huggingface.co/datasets/AIBench/M-VQA.beat-the-game-minecraft
Mine AI MCP — the run that beat Minecraft
An LLM agent played Minecraft 1.21.4 from an empty world to a defeated Ender Dragon,
autonomously, in a single unbroken session. No human input after the prompt, no
scripted behaviour trees, no save-scumming. This dataset is the complete record of
that run.
📺 Watch the run: https://www.youtube.com/watch?v=ZjtwWEfFVFY
💻 Code: https://github.com/aibengineering/mine-ai-mcp (MIT)
What it cost
Beating Minecraft took $93.17 of… See the full description on the dataset page: https://huggingface.co/datasets/aibengineering/beat-the-game-minecraft.ai-benchmarks-v2-2026
ai-benchmarks-v2-2026
AI data collected daily by Legion API.
🔑 API Access — Updated Daily
Live data via Legion AI API
Free: 100 req/day · Pro €29/month: 50K req/day + full fields
curl "https://api.legion-api.com/incidents?limit=10" -H "X-API-Key: YOUR_PRO_KEY"
Premium archive (1,285 incidents, full analysis): AISI Intelligence Pack €299
📦 Install
pip install legion-intel
from legion_intel import LegionClient
c = LegionClient()… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-benchmarks-v2-2026.EESE
The Ever-Evolving Science Exam (EESE)
AIBENCH
Dataset Description
Dataset Summary
As foundation models grow rapidly in capability and deployment, evaluating their scientific understanding becomes increasingly critical. Existing science benchmarks have made progress towards broad Range, wide Reach, and high Rigor, yet they often face two major challenges: data leakage risks that compromise benchmarking validity, and evaluation inefficiency due… See the full description on the dataset page: https://huggingface.co/datasets/AIBench/EESE.AIBE_mcq
Dataset Summary
This dataset is generated with the aim to collect all indian BAR exam questions. This could serve the purpose of Evaluating Language models.
Contributions
Mukesh Jha, DA-IICT, Gandhinagar, India
ai-benchmarks-2026
