datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RAG-Grounded-QA-188k
🎯 RAG Grounded QA 186K
The Anti-Hallucination Dataset
Teach language models to answer from context — or shut up trying.
Built by NovachronoAI — Precision AI for the real world.
Full Dataset (186K) · 20K Subset · Schema · Sources · Usage Guide
🧠 Why This Dataset Exists
Most QA datasets teach models what to say. This one also teaches them when to stay silent.
RAG (Retrieval-Augmented Generation) systems have a fatal flaw: the model hallucinates when… See the full description on the dataset page: https://huggingface.co/datasets/NovachronoAI/RAG-Grounded-QA-188k.Wikipedia_RAG_QA_Classification
🏛️ Wikipedia RAG QA Dataset for Retrieval-Augmented Generation Training
📊 Dataset Description
This dataset contains 300,000+ validated model-generated responses to Wikipedia content, specifically designed for Retrieval-Augmented Generation (RAG) applications and SQL database insertion tasks. Generated by Jeeney AI Reloaded 207M GPT with specialized RAG tuning.
🖥️ Demo Interface: Discord
Live Chat Demo on Discord: https://discord.gg/Xe9tHFCS9h
The full CJ… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Wikipedia_RAG_QA_Classification.ID_REG_MD_RAG
📑 Indonesian Regulation Markdown RAG Dataset (ID_REG_MD_RAG)
This repository contains a highly structured, Markdown-optimized collection of Indonesian Regulations (Peraturan Perundang-undangan). This dataset is specifically engineered to solve the "structure loss" problem often encountered when building Retrieval-Augmented Generation (RAG) systems for complex legal documents. 🏛️
💡 The Concept: Structural Integrity for RAG
Legal documents in Indonesia follow a… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/ID_REG_MD_RAG.RAGPulse
RAGPulse: A Real-World RAG Workload Trace to Optimize RAG Serving Systems
🌐 Github Link |
🤗 Workload Trace |
📑 Arxiv Paper |
🤖 How to use?
RAGPulse is a real-world RAG workload trace collected from an university-wide Q&A service scenario. The system has been serving over 40,000 students and faculties since April 2024, providing intelligent policy Q&A services. The trace contains a total of 7,106 records entries, sampled from one week of our Q&A service.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/flashserve/RAGPulse.elkarhizketak-RAG
Dataset Card for ElkarHizketak RAG and its Disruptor Variants
Base and disruptor variants of ElkarHizketak, built to stress-test conversational RAG systems in Basque under realistic interaction patterns (conversational openings, topic shifts).
Dataset Details
Dataset Description
This dataset extends ElkarHizketak with a base variant (rewritten opening queries, retrieval-needed labels, retrieved chunks) and disruptor variants that inject… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/elkarhizketak-RAG.arabic-rag-chat-8k-eval
arabic-rag-chat-8k-eval
Per-row evaluation artifacts for the 8,192-token Arabic multi-turn RAG models:
the test split, every model's raw replies, every judge verdict, and the rendered
report for each. Thirteen judged models, all scored on the same 1,651 prompts
by the same judge at temperature 0.0, so the comparison below is like-for-like
and can be recomputed offline without a GPU or a judge server.
This is the measurement half of
oddadmix/100M-8192-Nawah-dsv4;
the training… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-8k-eval.cs50-educational-rag
CS50 Pedagogical RAG Dataset
📜 Dataset Description
This repository contains the data artifacts for the undergraduate thesis, which explores the use of a pedagogical chatbot with Retrieval-Augmented Generation (RAG) for Harvard's CS50: Introduction to Computer Science course.
The project involved several stages of data processing, from raw content collection to the generation and curation of a high-quality evaluation dataset. To ensure full transparency and… See the full description on the dataset page: https://huggingface.co/datasets/dev-jonathanb/cs50-educational-rag.llm-rag-agent-papers
llm-rag-agent-papers
Research papers on LLM, RAG, and AI Agents - Knowledge base for RAG pipeline
Dataset Structure
This dataset contains three subsets:
llm: Large Language Model related content
rag: Retrieval-Augmented Generation related content
agent: AI Agent related content
Usage
from datasets import load_dataset
# Load all subsets
dataset = load_dataset("GXMZU/llm-rag-agent-papers")
# Load specific subset
llm_data =… See the full description on the dataset page: https://huggingface.co/datasets/GXMZU/llm-rag-agent-papers.arabic-rag-support-25K
Arabic RAG customer-support scenarios (27,927 rows)
Synthetic Modern Standard Arabic customer-support scenarios for training small
RAG answerers, distilled from unsloth/gemma-4-31B-it-NVFP4 on a local vLLM.
Built as the training set for oddadmix/Nawah-50M-RAG-Support.
Each row: a customer question + the knowledge-base chunks of one fictional
company (products, prices, policies, FAQ entries) + the ideal grounded agent
answer. One generation request invents one company KB and 4 QA… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-support-25K.rag-qa-arena
RAG QA Arena Annotated Dataset
A comprehensive multi-domain question-answering dataset with citation annotations designed for evaluating Retrieval-Augmented Generation (RAG) systems, featuring faithful answers with proper source attribution across 6 specialized domains.
🎯 Dataset Overview
This annotated version of the RAG QA Arena dataset includes citation information and gold document IDs, making it ideal for evaluating not just answer accuracy but also answer grounding… See the full description on the dataset page: https://huggingface.co/datasets/rajistics/rag-qa-arena.dog-rag
DOG-RAG: A Galician Benchmark for Legal Retrieval-Augmented Generation
Click to expand
Dataset description
Dataset Structure
Example
Entry Categories
Source Documents
Dataset Versions
Examples
Additional information
Acknowledgements
Cite this dataset
Dataset description
This dataset contains question–answer triplets derived from publications of the Diario Oficial de Galicia (DOG), the official gazette of the autonomous community of Galicia (Spain).… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/dog-rag.arabic-rag-chat-grpo-5K
Arabic multi-turn RAG conversations — GRPO pool (5,259 conversations)
The reinforcement-learning half of
oddadmix/arabic-rag-chat-30K:
same generator, same validator, same schema, disjoint companies. It exists
so GRPO explores fresh knowledge bases instead of taking a second pass over
material the SFT already memorised.
conversations
turns
companies
this pool
5,259
14,018
309
Company-disjointness is exact and verified: this pool shares zero
company_id values… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-grpo-5K.agmind-rag-splitter-ru-data
RU Context-Aware Document Split
Датасет (teacher-distillation) для обучения русского context-aware сплиттера документов для RAG. Каждый пример учит модель где резать документ на самодостаточные смысловые чанки, держа таблицы и код целыми.
Использован для модели AGmind/agmind-rag-splitter-ru. Код генерации и обучения: github.com/botAGI/AGmind-ML.
Формат (Alpaca JSONL)
{
"instruction": "Раздели документ на смысловые части для системы поиска (RAG)...",
"input":… See the full description on the dataset page: https://huggingface.co/datasets/AGmind/agmind-rag-splitter-ru-data.RAGPulse
RAGPulse: A Real-World RAG Workload Trace to Optimize RAG Serving Systems
🌐 Github Link |
🤗 Workload Trace |
📑 Arxiv Paper |
🤖 How to use?
RAGPulse is a real-world RAG workload trace collected from an university-wide Q&A service scenario. The system has been serving over 40,000 students and faculties since April 2024, providing intelligent policy Q&A services. The trace contains a total of 7,106 records entries, sampled from one week of our Q&A service.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/DiscoPosse/RAGPulse.multi-tafseer-quran-rag
Quran Tafseer RAG Dataset
A structured Arabic dataset of Quranic tafseer collected from eight classical and modern tafseer books.The dataset contains verse-aligned tafseer passages designed for Retrieval-Augmented Generation (RAG) systems and Arabic NLP research.
Each record links a Quran verse with its corresponding tafseer explanation from one of the tafseer books and includes rich metadata such as surah information, tafseer source, and embedding-ready text.
The dataset was… See the full description on the dataset page: https://huggingface.co/datasets/omaressam1111/multi-tafseer-quran-rag.aimi-anime-rag-dataset-sample
🎌 Ultimate Anime Dataset (8,248 Entries) | 1917-2025
A meticulously curated collection spanning 108 years of anime history
Love this dataset and the Anime Receipts concept? You can download the complete project via the links below:
🚀 Unlock the Full Potential
Product
What You Get
Get It Here
Tier 1
8,248 Anime Dataset (Parquet)
Tier 2
Full AiMi Recommendation System (Backend + UI)
Tier 3
Ultimate AiMi Recommendation System + AiMi Anime… See the full description on the dataset page: https://huggingface.co/datasets/DivyanshuSingh96/aimi-anime-rag-dataset-sample.enterprise-rag-internal-knowledge-search-benchmark-sample
Enterprise RAG and Internal Knowledge Search Benchmark Dataset -- Free Evaluation Sample
This dataset packages a synthetic internal company workspace and an evidence-linked benchmark table into one product for teams building enterprise RAG systems, internal search assistants, knowledge-base copilots, and deep-search evaluation pipelines.
The benchmark is designed around a realistic fictional company, AsteraOps Cloud, with multiple departments, renamed projects, stale roadmaps… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/enterprise-rag-internal-knowledge-search-benchmark-sample.arabic-rag-chat-30K
Arabic multi-turn RAG customer-support conversations (31,294 conversations)
Synthetic Modern Standard Arabic customer-support conversations for training
small Arabic RAG assistants. The bulk was distilled from gemini-3.1-flash-lite
via the Batch API; a first 2.4% came from unsloth/gemma-4-31B-it-NVFP4
on a local vLLM server before the run was moved off-GPU. Both teachers were given
the same prompts and the same validator. Each row is one conversation of 1-5
rounds over one… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-30K.Cosmopedia_QA_RAG_JSON_SQLiteThe full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com
Want more? 🚀 Get the AI Startup Bundle from Gumroad.
🖥️ Demo Interface: Discord
Discord: https://discord.gg/Xe9tHFCS9h
**Custom RAG QA generation services can be made available for paying customers to process internal documentation. DM me on Discord if you are interested.Jeeney AI GPT Reloaded 207M/Cosmopedia Model Outputs Dataset
Dataset Description
This dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Cosmopedia_QA_RAG_JSON_SQLite.fungi-rag-agent-sft-database
Fungi RAG Agent SFT Database
This dataset contains the supervised fine-tuning database used to train the
JacktheLander/smollm2-1.7b-fungi-rag-agent-distill-lora-gguf adapter.
It was built for the project-owned fungi RAG learning system:
JacktheLander/FunghiResearchAgent
The examples teach a small local language model to follow the system's agent contract:
use rag.search for evidence-backed mycology answers,
use safety.review for wild-mushroom edibility, field-identification… See the full description on the dataset page: https://huggingface.co/datasets/JacktheLander/fungi-rag-agent-sft-database.adaptive_rag_hotpotqa
Adaptive RAG HotpotQA Dataset
This dataset is a processed version of HotpotQA designed for training Adaptive Retrieval-Augmented Generation (RAG) systems.
Features
input: The input text for the model
output: The target output text
retrieval_label: Whether retrieval is needed (0/1)
hop: The reasoning hop number (1 or 2)
type: The type of example (multi_hop_qa, single_hop_qa, multi_hop_gating, etc.)
metadata: Additional information about the example including:
answer:… See the full description on the dataset page: https://huggingface.co/datasets/varun500/adaptive_rag_hotpotqa.enterprise-rag-internal-knowledge-search-benchmark
Enterprise RAG and Internal Knowledge Search Benchmark Dataset
This dataset packages a synthetic internal company workspace and an evidence-linked benchmark table into one product for teams building enterprise RAG systems, internal search assistants, knowledge-base copilots, and deep-search evaluation pipelines.
The benchmark is designed around a realistic fictional company, AsteraOps Cloud, with multiple departments, renamed projects, stale roadmaps, support escalations… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/enterprise-rag-internal-knowledge-search-benchmark.
