datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pdfQA-Annotations
pdfQA: Diverse, Challenging, and Realistic Question Answering over PDFs
pdfQA is a structured benchmark collection for document-level question answering and PDF understanding research.
This repository contains the pdfQA-Annotations dataset, which provides only the QA annotations and metadata for the pdfQA-Benchmark.
It is intended for lightweight experimentation, modeling, and evaluation without requiring access to large document files.
Relationship to the Full pdfQA… See the full description on the dataset page: https://huggingface.co/datasets/pdfqa/pdfQA-Annotations.wearable-agent-trajectory-annotations
Wearable Agent Trajectory Annotation Dataset
Dataset Summary
50 wearable agent trajectories annotated by 5 LLM-simulated annotator personas
using the agenteval-schema-v1 JSON schema, across two calibration phases
(500 annotation records total). Designed to benchmark annotation-quality pipelines
for agentic AI systems.
Each trajectory captures a wearable AI agent responding to a real-time sensor event
(health alert, privacy-sensitive context, location trigger… See the full description on the dataset page: https://huggingface.co/datasets/finaspirant/wearable-agent-trajectory-annotations.crowdsourced-vru-annotations
Crowdsourced VRU Annotations
Dataset Summary
This dataset provides tabular annotations from two underlying datasets — ECP and ZOD — and is organized into two splits (ecp and zod).It contains the results of crowdsourced annotation tasks focusing on vulnerable road users (VRUs).
The underlying examples are image crops showing bounding boxes of VRUs from the ECP and ZOD datasets. Each crop is referenced via a crop_id. The actual image pixels are not included in this… See the full description on the dataset page: https://huggingface.co/datasets/cklugmann/crowdsourced-vru-annotations.DAM-QA-annotations
DAM-QA Unified Annotations
22,675 question-answer pairs from 6 major VQA benchmarks, unified for the DAM-QA framework. This collection consolidates annotations from InfographicVQA, TextVQA, VQAv2, DocVQA, ChartQA, and ChartQA-Pro into standardized JSONL formats.
📖 Paper: Describe Anything Model for Visual Question Answering on Text-rich Images⚠️ Note: Images not included - obtain from original sources with proper licensing
Repository Structure
DAM-QA-annotations/… See the full description on the dataset page: https://huggingface.co/datasets/VLAI-AIVN/DAM-QA-annotations.llm-delusion-response-annotations
LLM Delusion-Like Belief Reinforcement Annotations
This dataset contains human annotations of responses generated by conversational large language models (LLMs) to prompts expressing potentially delusion-like or reality-distorted beliefs.
The purpose of the dataset is to support evaluation of whether conversational LLM responses may unintentionally reinforce or strengthen delusion-like beliefs.
Dataset Files
Consensus Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ManjuKrish/llm-delusion-response-annotations.llm-delusion-response-annotations
LLM Delusion-Like Belief Reinforcement Annotations
This dataset contains human annotations of responses generated by conversational large language models (LLMs) to prompts expressing potentially delusion-like or reality-distorted beliefs.
The purpose of the dataset is to support evaluation of whether conversational LLM responses may unintentionally reinforce or strengthen delusion-like beliefs.
Dataset Files
Consensus Dataset… See the full description on the dataset page: https://huggingface.co/datasets/vennu95/llm-delusion-response-annotations.TruthQuest-AI-AnnotationsLiar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models
This data repository contains the model answers and LLM-based (conclusion and error) annotations from the paper Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models (Mondorf and Plank, 2024).
Below, we provide a short description of each column in our dataset:
Statement Set (Literal["S", "I", "E"]): The type of statement set used in the puzzle.
Problem (list of… See the full description on the dataset page: https://huggingface.co/datasets/mainlp/TruthQuest-AI-Annotations.TruthQuest-Human-AnnotationsLiar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models
This data repository contains the model answers and human (conclusion and error) annotations from the paper Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models (Mondorf and Plank, 2024).
Below, we provide a short description of each column in our dataset:
Statement Set (Literal["S", "I", "E"]): The type of statement set used in the puzzle.
Problem (list of… See the full description on the dataset page: https://huggingface.co/datasets/mainlp/TruthQuest-Human-Annotations.
