datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HumaniBench
HumaniBench: A Human-Centric Benchmark for Large Multimodal Models Evaluation
**HumaniBench** is a benchmark for evaluating large multimodal models (LMMs) using real-world, human-centric criteria. It consists of 32,000+ image–question pairs across 7 tasks:
✅ Open/closed VQA
🌍 Multilingual QA
📌 Visual grounding
💬 Empathetic captioning
🧠 Robustness, reasoning, and ethics
Each example is annotated with GPT-4o drafts, then verified by experts to ensure quality and… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/HumaniBench.arXiv-AI-papers-multi-vector
Overview
This is a dataset containing individual pages from the top-40 most cited AI papers on arXiv](https://arxiv.org/abs/2412.12121) from the period 2023-01-01 to 2024-09-30.
Only the first 10 pages from each paper is included.
The dataset includes an image of each page as well as a multi-vector embedding using vidore/colqwen2-v1.0.
ibm-hls-burn-vectorized
