agentic-rag
AgenticRAGTracer
AgenticRAGTracer: A Hop-Aware Benchmark for Diagnosing Multi-Step Retrieval Reasoning in Agentic RAG
Paper | Code
🎉 Our work has been accepted to ACL 2026 Findings!
AgenticRAGTracer is a benchmark designed to diagnose and evaluate multi-step retrieval reasoning in Agentic RAG systems. Unlike traditional benchmarks that provide only final questions and answers, AgenticRAGTracer includes intermediate hop-level questions that connect atomic questions to the final query. This allows… See the full description on the dataset page: https://huggingface.co/datasets/YqjMartin/AgenticRAGTracer.agentic-rag-redteam-bench
WARNING: HARMFUL CONTENT - RESEARCH USE ONLY
This dataset contains adversarial prompts, jailbreak attacks, toxic outputs, and other explicitly harmful content generated for AI safety research. Samples include prompt injections, social engineering payloads, misinformation, hate speech, instructions for illegal activities, phishing templates, and other dangerous material. All content is synthetic and produced by automated red-teaming pipelines for the sole purpose of evaluating and improving… See the full description on the dataset page: https://huggingface.co/datasets/Fujitsu/agentic-rag-redteam-bench.secai-agentic-rag-sft-v1
secAI Agentic-RAG SFT v1
A curated Vietnamese/English cybersecurity instruction and agentic-RAG supervised fine-tuning dataset. It teaches direct security assistance as well as grounded tool-use behaviour: tool selection, JSON arguments, consuming tool results, no-result handling, and multi-turn follow-ups.
This is a frozen training release, not a standalone claim of safety, factual correctness, or production readiness. Keep human review and authorization controls for all… See the full description on the dataset page: https://huggingface.co/datasets/DuyTa/secai-agentic-rag-sft-v1.AgenticRag
AgenticRAG-FP
This repository describes AgenticRAG-FP, a research dataset and evaluation
suite for studying how failures propagate through agentic retrieval-augmented
generation pipelines. The dataset normalizes multi-hop QA examples into a common
schema, runs real or mock ReAct-style RAG agents over them, injects controlled
failures at specific retrieval/reasoning hops, and records whether diagnostic
methods can recover the true root cause after the failure has propagated.
The… See the full description on the dataset page: https://huggingface.co/datasets/anote-ai/AgenticRag.dataflow-demo-multihop-AgenticRAG
DataFlow Demo -- AgenticRAG Pipeline
This dataset card serves as a demo for showcasing the Agentic Retrieval-Augmented Generation (AgenticRAG) pipeline of the DataFlow Project.It provides a clear comparison between single-hop input questions and their corresponding multi-hop, agent-driven reasoning outputs.
Overview
The AgenticRAG pipeline is designed to solve complex questions that require multi-step reasoning across multiple documents.Instead of answering questions in… See the full description on the dataset page: https://huggingface.co/datasets/OpenDCAI/dataflow-demo-multihop-AgenticRAG.alpha_agentic_search_ragdata
