datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AgenticRAGTracer
AgenticRAGTracer: A Hop-Aware Benchmark for Diagnosing Multi-Step Retrieval Reasoning in Agentic RAG
Paper | Code
🎉 Our work has been accepted to ACL 2026 Findings!
AgenticRAGTracer is a benchmark designed to diagnose and evaluate multi-step retrieval reasoning in Agentic RAG systems. Unlike traditional benchmarks that provide only final questions and answers, AgenticRAGTracer includes intermediate hop-level questions that connect atomic questions to the final query. This allows… See the full description on the dataset page: https://huggingface.co/datasets/YqjMartin/AgenticRAGTracer.agentic-rag-redteam-bench
WARNING: HARMFUL CONTENT - RESEARCH USE ONLY
This dataset contains adversarial prompts, jailbreak attacks, toxic outputs, and other explicitly harmful content generated for AI safety research. Samples include prompt injections, social engineering payloads, misinformation, hate speech, instructions for illegal activities, phishing templates, and other dangerous material. All content is synthetic and produced by automated red-teaming pipelines for the sole purpose of evaluating and improving… See the full description on the dataset page: https://huggingface.co/datasets/Fujitsu/agentic-rag-redteam-bench.AgenticRag
AgenticRAG-FP
This repository describes AgenticRAG-FP, a research dataset and evaluation
suite for studying how failures propagate through agentic retrieval-augmented
generation pipelines. The dataset normalizes multi-hop QA examples into a common
schema, runs real or mock ReAct-style RAG agents over them, injects controlled
failures at specific retrieval/reasoning hops, and records whether diagnostic
methods can recover the true root cause after the failure has propagated.
The… See the full description on the dataset page: https://huggingface.co/datasets/anote-ai/AgenticRag.agentic-rag-redteam-bench
Mirror note: This dataset is a mirror of Fujitsu/agentic-rag-redteam-bench, maintained by the same author, intherejeet, to preserve availability if organization access is interrupted. Access controls and usage restrictions are intended to match the source dataset.
WARNING: HARMFUL CONTENT - RESEARCH USE ONLY
This dataset contains adversarial prompts, jailbreak attacks, toxic outputs, and other explicitly harmful content generated for AI safety research. Samples include prompt injections… See the full description on the dataset page: https://huggingface.co/datasets/intherejeet/agentic-rag-redteam-bench.
