datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AgenticRAGTracer
AgenticRAGTracer: A Hop-Aware Benchmark for Diagnosing Multi-Step Retrieval Reasoning in Agentic RAG
Paper | Code
🎉 Our work has been accepted to ACL 2026 Findings!
AgenticRAGTracer is a benchmark designed to diagnose and evaluate multi-step retrieval reasoning in Agentic RAG systems. Unlike traditional benchmarks that provide only final questions and answers, AgenticRAGTracer includes intermediate hop-level questions that connect atomic questions to the final query. This allows… See the full description on the dataset page: https://huggingface.co/datasets/YqjMartin/AgenticRAGTracer.AgenticRag
AgenticRAG-FP
This repository describes AgenticRAG-FP, a research dataset and evaluation
suite for studying how failures propagate through agentic retrieval-augmented
generation pipelines. The dataset normalizes multi-hop QA examples into a common
schema, runs real or mock ReAct-style RAG agents over them, injects controlled
failures at specific retrieval/reasoning hops, and records whether diagnostic
methods can recover the true root cause after the failure has propagated.
The… See the full description on the dataset page: https://huggingface.co/datasets/anote-ai/AgenticRag.
