datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rag_hallucinationsProvides examples of hallucinated responses for RAG applications.
simutrade-rag-sft-28k
📢 Domain & Email Migration Notice
From May 6th, 2026, Simutrade will transition to new domains as simutrade.app will not be renewed:
🌐 Website: simutrade.faizath.com (formerly simutrade.app)
⚙️ API: simutrade-api.faizath.com (formerly api.simutrade.app)
📧 Email: contact@simutrade.faizath.com (formerly contact@simutrade.app)
🛰️ CDN: simutrade-cdn.faizath.com (formerly cdn.simutrade.app)
📈 Status Pages:… See the full description on the dataset page: https://huggingface.co/datasets/simutrade/simutrade-rag-sft-28k.RAG-Instruct
Introduction
RAG-Instruct is a RAG dataset designed to comprehensively enhance LLM RAG capabilities, synthesized using GPT-4o. This dataset is based on the Wikipedia corpus and This dataset is based on the Wikipedia corpus and offers the advantages of query-document scenario diversity and task diversity.
The RAG-Instruct dataset can significantly enhance the RAG ability of LLMs and make remarkable improvements in RAG performance across various tasks.
Model
WQA (acc)
PQA (acc)… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/RAG-Instruct.RAG_Multilingual
Dataset Card for RAG_Multilingual
Dataset Summary
RAG_Multilingual is an instruction-following synthetic QA dataset created from extractive QA datasets from Catalan, English and Spanish reference sets.
The reference datasets were: SQAD (https://huggingface.co/datasets/rajpurkar/squad), Catalanqa (https://huggingface.co/datasets/projecte-aina/catalanqa) and SQAC (https://huggingface.co/datasets/PlanTL-GOB-ES/SQAC).
This dataset, of 56.406 instances, was created by… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/RAG_Multilingual.agentic-rag-redteam-bench
WARNING: HARMFUL CONTENT - RESEARCH USE ONLY
This dataset contains adversarial prompts, jailbreak attacks, toxic outputs, and other explicitly harmful content generated for AI safety research. Samples include prompt injections, social engineering payloads, misinformation, hate speech, instructions for illegal activities, phishing templates, and other dangerous material. All content is synthetic and produced by automated red-teaming pipelines for the sole purpose of evaluating and improving… See the full description on the dataset page: https://huggingface.co/datasets/Fujitsu/agentic-rag-redteam-bench.cfr-rag-jsonECFR from 06/2025
Wikipedia_RAG_QA_Classification
🏛️ Wikipedia RAG QA Dataset for Retrieval-Augmented Generation Training
📊 Dataset Description
This dataset contains 300,000+ validated model-generated responses to Wikipedia content, specifically designed for Retrieval-Augmented Generation (RAG) applications and SQL database insertion tasks. Generated by Jeeney AI Reloaded 207M GPT with specialized RAG tuning.
🖥️ Demo Interface: Discord
Live Chat Demo on Discord: https://discord.gg/Xe9tHFCS9h
The full CJ… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Wikipedia_RAG_QA_Classification.LUNA-RAG-MCP-SFT-10M
Dataset Card for LUNA RAG + MCP SFT Dataset
A clean, English-only, instruction-finetuning dataset for teaching small language models two of the most important 2025–2026 agentic-AI topics: Retrieval-Augmented Generation (RAG) and the Model Context Protocol (MCP).
This repository is the instruction-tuning (SFT) companion dataset for the LUNA-100M model family. It is intentionally compact (≈10M formatted tokens, ≤1,024 tokens per sample) so that it can be absorbed efficiently by a… See the full description on the dataset page: https://huggingface.co/datasets/ASTERIZER/LUNA-RAG-MCP-SFT-10M.Finetune-RAG
Finetune-RAG Dataset
This dataset is part of the Finetune-RAG project, which aims to tackle hallucination in retrieval-augmented LLMs. It consists of synthetically curated and processed RAG documents that can be utilised for LLM fine-tuning.
Each line in the finetunerag_dataset.jsonl file is a JSON object:
{
"content": "<correct content chunk retrieved>",
"filename": "<original document filename>",
"fictitious_filename1":"<filename of fake doc 1>",
"fictitious_content1":… See the full description on the dataset page: https://huggingface.co/datasets/pints-ai/Finetune-RAG.humanevalHumanEval dataset annotated with the ground-truth programming solutions, to enable evaluations for retrieval and retrieval augmented code generation.
Please refer to code-rag-becnch for more details.
LIT-RAGBench
LIT-RAGBench
LIT-RAGBench is a benchmark for evaluating generator capabilities in Retrieval-Augmented Generation (RAG). It focuses on whether a model can answer questions correctly given retrieved documents, independent of retrieval quality. The benchmark covers five categories: Integration, Reasoning, Logic, Table, and Abstention.
Dataset Summary
LIT-RAGBench contains:
114 human-constructed Japanese questions
An English version generated by machine translation with… See the full description on the dataset page: https://huggingface.co/datasets/neoai-inc/LIT-RAGBench.rage-core-1
Rage Core 1
Supervised fine-tuning dataset for Rage Core Gen 1 — developed by DevNameGelo, Powered by RGC Rage Gen Core.
This dataset teaches the model to: understand user intent and ask clarifying questions when requirements are ambiguous, deliver complete production-oriented implementations, debug real-world problems, make precise code edits, operate as a coding agent through explicit tool calls, handle multilingual users, process multimodal inputs honestly, and refuse to… See the full description on the dataset page: https://huggingface.co/datasets/rgcmainhub/rage-core-1.RAGPulse
RAGPulse: A Real-World RAG Workload Trace to Optimize RAG Serving Systems
🌐 Github Link |
🤗 Workload Trace |
📑 Arxiv Paper |
🤖 How to use?
RAGPulse is a real-world RAG workload trace collected from an university-wide Q&A service scenario. The system has been serving over 40,000 students and faculties since April 2024, providing intelligent policy Q&A services. The trace contains a total of 7,106 records entries, sampled from one week of our Q&A service.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/flashserve/RAGPulse.mbppMBPP dataset annotated with ground-truth programming solutions, to enable evaluations for retrieval and retrieval-augmented code generation.
Please refer to code-rag-bench for more details.
elkarhizketak-RAG
Dataset Card for ElkarHizketak RAG and its Disruptor Variants
Base and disruptor variants of ElkarHizketak, built to stress-test conversational RAG systems in Basque under realistic interaction patterns (conversational openings, topic shifts).
Dataset Details
Dataset Description
This dataset extends ElkarHizketak with a base variant (rewritten opening queries, retrieval-needed labels, retrieved chunks) and disruptor variants that inject… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/elkarhizketak-RAG.arabic-rag-chat-8k-eval
arabic-rag-chat-8k-eval
Per-row evaluation artifacts for the 8,192-token Arabic multi-turn RAG models:
the test split, every model's raw replies, every judge verdict, and the rendered
report for each. Thirteen judged models, all scored on the same 1,651 prompts
by the same judge at temperature 0.0, so the comparison below is like-for-like
and can be recomputed offline without a GPU or a judge server.
This is the measurement half of
oddadmix/100M-8192-Nawah-dsv4;
the training… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-8k-eval.odexODEX dataset annotated with the ground-truth library documentation, to enable evaluations for retrieval and retrieval-augmented code generation.
Please refer to [code-rag-bench] for more details.
Hiro-Pharma-RAG-Benchmark
Hiro Pharma RAG Benchmark
This private dataset repository contains multilingual biomedical RAG benchmark data associated with the paper CRAB: A Benchmark for Evaluating Curation of Retrieval-Augmented LLMs in Biomedicine.
The benchmark is designed to evaluate whether retrieval-augmented language models can answer biomedical questions while selecting and citing useful evidence and filtering out noisy or irrelevant references.
Repository Contents
File
Language… See the full description on the dataset page: https://huggingface.co/datasets/PatSnap/Hiro-Pharma-RAG-Benchmark.wikipedia-rag
Wikipedia RAG
This dataset consists of raw wikipedia pages with one question and answer for each page. The answers were generated using Gemma 3-27B.
This dataset is provided as part of the AMALIA project and is included in the data mix used to post-train the AMALIA model.
Citation
If you use this dataset or AMALIA in your work, please cite:
@inproceedings{simplicio-etal-2026-amalia,
title = "{AMALIA}: A Fully Open Large Language Model for… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/wikipedia-rag.NQ-RAG-DPO-Evaluation
Dataset Card
Dataset Summary
This repository contains five structured data files forming a complete evaluation and training workflow for a multi-perspective alignment system built on Retrieval-Augmented Generation (RAG) and Direct Preference Optimization (DPO).
The system is organized into three interconnected pipelines:
1️. RAG Pipeline
The RAG pipeline uses the Natural Questions (validation split) as the knowledge source and evaluation benchmark.
For each… See the full description on the dataset page: https://huggingface.co/datasets/AnjanSB/NQ-RAG-DPO-Evaluation.movie-screenplays-tokenized-dataset
Screenplay Corpus — Tokenized (GPT-2)
Pre-tokenized screenplay corpus used to train and evaluate the models in the GPT-2 Screenplay Fine-Tuning Study. Derived from Movie-Script-Database by Aveek Saha. Provided as tokenized JSON splits ready for direct consumption by a GPT-2 Trainer pipeline — no preprocessing required.
Dataset Description
This dataset contains approximately 94 million tokens of professionally formatted screenplay text, pre-tokenized using the… See the full description on the dataset page: https://huggingface.co/datasets/raghavnimbalkar/movie-screenplays-tokenized-dataset.rag-systems-sft-100k
RAG Systems SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality conversations covering Retrieval-Augmented Generation (RAG) systems — from basic pipelines to advanced multi-hop retrieval, evaluation, and production optimization. Designed to train AI assistants that can help engineers build, debug, and scale RAG applications.
Dataset Description
This dataset covers the full spectrum of RAG system development across 12 specialized categories.… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/rag-systems-sft-100k.arabic-rag-support-25K
Arabic RAG customer-support scenarios (27,927 rows)
Synthetic Modern Standard Arabic customer-support scenarios for training small
RAG answerers, distilled from unsloth/gemma-4-31B-it-NVFP4 on a local vLLM.
Built as the training set for oddadmix/Nawah-50M-RAG-Support.
Each row: a customer question + the knowledge-base chunks of one fictional
company (products, prices, policies, FAQ entries) + the ideal grounded agent
answer. One generation request invents one company KB and 4 QA… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-support-25K.ml-interview-sft-dataset
ML/AI Interview Coach — SFT Dataset
A curated dataset of 566 high-quality Q&A pairs covering ML, Deep Learning, NLP, LLMs, RAG, Vector Databases, LangChain, Agentic AI, MLOps, and more — designed for fine-tuning an ML Interview Coach model.
Dataset Summary
Stat
Value
Total Q&A pairs
566
Unique topics
75
Format
ChatML (system + user + assistant)
Language
English
Avg answer length
~800 tokens
Sources
15+ interview prep documents + hand-crafted… See the full description on the dataset page: https://huggingface.co/datasets/raghu298/ml-interview-sft-dataset.drtulu_v2_personalfinance
Personal Finance Reasoning-V2.1
This dataset is associated with the paper Synthesizing Behaviorally-Grounded Reasoning Chains: A Data-Generation Framework for Personal Finance LLMs.
This is a scaled up version of the PersonalFinance-V2 dataset with some pipeline streamlining done.*
1. Introduction & Motivation
The landscape of financial AI benchmarks is currently dominated by applications in corporate finance, algorithmic trading, and general financial knowledge… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/drtulu_v2_personalfinance.zarn-workspace-rag-qa
Zarn Workspace RAG QA
Dataset Description
Small document bundles paired with grounded answers, evidence, and explicit refusals when context is missing.
Team Attribution
This dataset was created and reviewed by the Zarnite team through internal benchmark design, generation, and quality-control workflows. It should be presented as a Zarnite-authored benchmark starter pack, not as a purely human-collected field corpus.
Ecosystem Need Tier
High Ecosystem… See the full description on the dataset page: https://huggingface.co/datasets/zarnite/zarn-workspace-rag-qa.CodeSecAudit-RAG
CodeSecAudit-RAG
CodeSecAudit-RAG is a curated defensive dataset for building an Enterprise Code Review and Security Auditor Agent. It combines vulnerability-detection examples with a retrieval-ready secure-coding knowledge corpus.
The dataset is designed for practical AIML and MLOps workflows such as vulnerability detection, security review explanation, and RAG-based secure coding guidance retrieval.
Dataset Components
1. Review Dataset
Files:… See the full description on the dataset page: https://huggingface.co/datasets/OMCHOKSI108/CodeSecAudit-RAG.CoT-Scientific-RAG-Reasoning
CoT-Scientific-RAG-Reasoning
This dataset is designed for fine-tuning Large Language Models (specifically Qwen-series) to perform complex reasoning over scientific and technical documents using Chain-of-Thought (CoT).
Dataset Description
The dataset contains instructions and scientific contexts (Medical Imaging, Autonomous Driving, VLA Frameworks) where the model is required to generate a reasoning trace before providing the final answer.
Format: JSONL
Logic: All outputs… See the full description on the dataset page: https://huggingface.co/datasets/abhinavdread/CoT-Scientific-RAG-Reasoning.dog-rag
DOG-RAG: A Galician Benchmark for Legal Retrieval-Augmented Generation
Click to expand
Dataset description
Dataset Structure
Example
Entry Categories
Source Documents
Dataset Versions
Examples
Additional information
Acknowledgements
Cite this dataset
Dataset description
This dataset contains question–answer triplets derived from publications of the Diario Oficial de Galicia (DOG), the official gazette of the autonomous community of Galicia (Spain).… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/dog-rag.ds1000DS-1000 dataset annotated with the ground-truth library documentation, to enable evaluations for retrieval and retrieval-augmented code generation.
Please refer to [code-rag-bench] for more details
