CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01aporia-ai /rag_hallucinationsProvides examples of hallucinated responses for RAG applications. textquestion-answering1K<n<10K9 likes462 downloads2y agoHugging Face02ai4bharat /Indic-Rag-Suite 🌏 Multilingual Indic RAG Suite A comprehensive multilingual question-answering dataset covering 18 Indian languages with 21,439,886 total samples, designed for RAG (Retrieval-Augmented Generation) applications and multilingual NLP research. 🚀 Quick Start from datasets import load_dataset # Load specific language (recommended) dataset = load_dataset("ai4bharat/Indic-Rag-Suite", "as") train_data = dataset['train'] print(f"Loaded {len(train_data)} samples") # Access… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/Indic-Rag-Suite.textquestion-answering10M<n<100M2 likes305 downloads1y agoHugging Face03airesearch /WangchanX-Legal-ThaiCCL-RAG 🏛️ WangchanX-Legal-ThaiCCL-RAG [Technical Report] The WangchanX-Legal-ThaiCCL-RAG dataset supports the development Retrieval-Augmented Generation (RAG) for Thai Legal question answering. This dataset is allows developers to finetune both retrieval model - to better retrieve relevant law section, and Large Language Model (LLM) - for instruction tuning. Our dataset supports Corporate and Commercial Law (thus ThaiCCL name). See legislation section for more details on supported… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/WangchanX-Legal-ThaiCCL-RAG.texttext-generation10K<n<100K13 likes283 downloads2y agoHugging Face04simutrade /simutrade-rag-sft-28k 📢 Domain & Email Migration Notice From May 6th, 2026, Simutrade will transition to new domains as simutrade.app will not be renewed: 🌐 Website: simutrade.faizath.com (formerly simutrade.app) ⚙️ API: simutrade-api.faizath.com (formerly api.simutrade.app) 📧 Email: contact@simutrade.faizath.com (formerly contact@simutrade.app) 🛰️ CDN: simutrade-cdn.faizath.com (formerly cdn.simutrade.app) 📈 Status Pages:… See the full description on the dataset page: https://huggingface.co/datasets/simutrade/simutrade-rag-sft-28k.textquestion-answering10K<n<100K1 likes216 downloads1mo agoHugging Face05FreedomIntelligence /RAG-Instruct Introduction RAG-Instruct is a RAG dataset designed to comprehensively enhance LLM RAG capabilities, synthesized using GPT-4o. This dataset is based on the Wikipedia corpus and This dataset is based on the Wikipedia corpus and offers the advantages of query-document scenario diversity and task diversity. The RAG-Instruct dataset can significantly enhance the RAG ability of LLMs and make remarkable improvements in RAG performance across various tasks. Model WQA (acc) PQA (acc)… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/RAG-Instruct.textquestion-answering10K<n<100K43 likes185 downloads2y agoHugging Face06projecte-aina /RAG_Multilingual Dataset Card for RAG_Multilingual Dataset Summary RAG_Multilingual is an instruction-following synthetic QA dataset created from extractive QA datasets from Catalan, English and Spanish reference sets. The reference datasets were: SQAD (https://huggingface.co/datasets/rajpurkar/squad), Catalanqa (https://huggingface.co/datasets/projecte-aina/catalanqa) and SQAC (https://huggingface.co/datasets/PlanTL-GOB-ES/SQAC). This dataset, of 56.406 instances, was created by… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/RAG_Multilingual.textquestion-answering10K<n<100K23 likes176 downloads2y agoHugging Face07NovachronoAI /RAG-Grounded-QA-188k 🎯 RAG Grounded QA 186K The Anti-Hallucination Dataset Teach language models to answer from context — or shut up trying. Built by NovachronoAI — Precision AI for the real world. Full Dataset (186K) · 20K Subset · Schema · Sources · Usage Guide 🧠 Why This Dataset Exists Most QA datasets teach models what to say. This one also teaches them when to stay silent. RAG (Retrieval-Augmented Generation) systems have a fatal flaw: the model hallucinates when… See the full description on the dataset page: https://huggingface.co/datasets/NovachronoAI/RAG-Grounded-QA-188k.tabularquestion-answering100K<n<1M0 likes167 downloads7mo agoHugging Face08Fujitsu /agentic-rag-redteam-benchgated WARNING: HARMFUL CONTENT - RESEARCH USE ONLY This dataset contains adversarial prompts, jailbreak attacks, toxic outputs, and other explicitly harmful content generated for AI safety research. Samples include prompt injections, social engineering payloads, misinformation, hate speech, instructions for illegal activities, phishing templates, and other dangerous material. All content is synthetic and produced by automated red-teaming pipelines for the sole purpose of evaluating and improving… See the full description on the dataset page: https://huggingface.co/datasets/Fujitsu/agentic-rag-redteam-bench.imagetext-retrieval10K<n<100K1 likes165 downloads7mo agoHugging Face09MedinaArmando /cfr-rag-jsonECFR from 06/2025 texttext-generation100K<n<1M0 likes161 downloads1y agoHugging Face10hoskerelab /ragalyst-qac RAGalyst-QAC RAGalyst-QAC dataset is a collection of synthetically generated domain-specific question-answer-context (QAC) triplets designed to evaluate retrieval-augmented generation (RAG) systems. We provide 500 QAC triplets across three impacful domains: military operations, cybersecurity, and bridge engineering. Dataset Sources Repository: Github Paper: Arxiv Website: RAGalyst Pip Package: Coming soon! Dataset Structure Each sample is a QAC triplet… See the full description on the dataset page: https://huggingface.co/datasets/hoskerelab/ragalyst-qac.documentquestion-answering10K<n<100K0 likes128 downloads10mo agoHugging Face11f20180301 /loft-rag-nq-128k LOFT RAG - Natural Questions (128k) Dataset Description This dataset is part of the LOFT (Long-context Open Foundation Tasks) benchmark, specifically the RAG (Retrieval-Augmented Generation) task. Dataset: Natural Questions Context Length: 128k Task Type: RAG (Retrieval-Augmented Generation) Language: English Source: LOFT Benchmark (Google DeepMind) Dataset Structure Data Fields context (string): Full prompt context including corpus… See the full description on the dataset page: https://huggingface.co/datasets/f20180301/loft-rag-nq-128k.textquestion-answeringn<1K0 likes116 downloads10mo agoHugging Face12CJJones /Wikipedia_RAG_QA_Classification 🏛️ Wikipedia RAG QA Dataset for Retrieval-Augmented Generation Training 📊 Dataset Description This dataset contains 300,000+ validated model-generated responses to Wikipedia content, specifically designed for Retrieval-Augmented Generation (RAG) applications and SQL database insertion tasks. Generated by Jeeney AI Reloaded 207M GPT with specialized RAG tuning. 🖥️ Demo Interface: Discord Live Chat Demo on Discord: https://discord.gg/Xe9tHFCS9h The full CJ… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Wikipedia_RAG_QA_Classification.tabulartext-generation100K<n<1M1 likes112 downloads6mo agoHugging Face13jakobsnel /RAGTruth_Xtended Dataset Card for Dataset Name This dataset provides response token logits and hidden states, complementing the underlying RAGTruth dataset. It has been generated using https://github.com/jakobsnl/RAGTruth_Xtended. Dataset Details Dataset Description This dataset is built upon RAGTruth (github.com/ParticleMedia/RAGTruth), which consists of character-level annotation of different types of hallucination for responses to a given set of LLM tasks. Out of all models… See the full description on the dataset page: https://huggingface.co/datasets/jakobsnel/RAGTruth_Xtended.texttext-generation10K<n<100K0 likes108 downloads1y agoHugging Face14ASTERIZER /LUNA-RAG-MCP-SFT-10M Dataset Card for LUNA RAG + MCP SFT Dataset A clean, English-only, instruction-finetuning dataset for teaching small language models two of the most important 2025–2026 agentic-AI topics: Retrieval-Augmented Generation (RAG) and the Model Context Protocol (MCP). This repository is the instruction-tuning (SFT) companion dataset for the LUNA-100M model family. It is intentionally compact (≈10M formatted tokens, ≤1,024 tokens per sample) so that it can be absorbed efficiently by a… See the full description on the dataset page: https://huggingface.co/datasets/ASTERIZER/LUNA-RAG-MCP-SFT-10M.texttext-generation10K<n<100K0 likes108 downloads27d agoHugging Face15pints-ai /Finetune-RAG Finetune-RAG Dataset This dataset is part of the Finetune-RAG project, which aims to tackle hallucination in retrieval-augmented LLMs. It consists of synthetically curated and processed RAG documents that can be utilised for LLM fine-tuning. Each line in the finetunerag_dataset.jsonl file is a JSON object: { "content": "<correct content chunk retrieved>", "filename": "<original document filename>", "fictitious_filename1":"<filename of fake doc 1>", "fictitious_content1":… See the full description on the dataset page: https://huggingface.co/datasets/pints-ai/Finetune-RAG.texttext-generation1K<n<10K6 likes104 downloads1y agoHugging Face16code-rag-bench /humanevalHumanEval dataset annotated with the ground-truth programming solutions, to enable evaluations for retrieval and retrieval augmented code generation. Please refer to code-rag-becnch for more details. texttext-generationn<1K0 likes99 downloads2y agoHugging Face17neoai-inc /LIT-RAGBench LIT-RAGBench LIT-RAGBench is a benchmark for evaluating generator capabilities in Retrieval-Augmented Generation (RAG). It focuses on whether a model can answer questions correctly given retrieved documents, independent of retrieval quality. The benchmark covers five categories: Integration, Reasoning, Logic, Table, and Abstention. Dataset Summary LIT-RAGBench contains: 114 human-constructed Japanese questions An English version generated by machine translation with… See the full description on the dataset page: https://huggingface.co/datasets/neoai-inc/LIT-RAGBench.textquestion-answeringn<1K0 likes99 downloads5mo agoHugging Face18OpenLLM-France /Luciole_RAG Dataset overview Luciole RAG is a supervised fine-tuning dataset for retrieval-augmented generation, built to train the Luciole models. Each example is a chat conversation where the assistant answers a question using only a set of retrieved document chunks given in the system prompt, quotes and cites its sources, and declines to answer when the documents do not contain the answer. It contains two subsets derived from existing question-answering benchmarks: Config Source… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole_RAG.textquestion-answering10K<n<100K1 likes96 downloads2d agoHugging Face19sdiazlor /rag-human-rights-from-files Dataset Card for my-distiset-rag-files This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/my-distiset-rag-files/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/rag-human-rights-from-files.texttext-generationn<1K0 likes90 downloads2y agoHugging Face20cnmoro /RagMixPTBR-Legal-Alpaca-2M Este é um dataset que é composto por 2 datasets menores: cnmoro/WizardVicuna-PTBR-Instruct-Clean cnmoro/GPT4-500k-Augmented-PTBR-Clean Além destes dois, foi desenvolvido um novo dataset criado sinteticamente, que utiliza do formato “Alpaca”, contendo não apenas duas, mas três informações segmentadas: Contexto/Input Pergunta Resposta Para a criação desse bloco, foi utilizado o eduagarcia/LegalPT_dedup como base, objetivando incorporar informações na área do direito (além de outros datasets… See the full description on the dataset page: https://huggingface.co/datasets/cnmoro/RagMixPTBR-Legal-Alpaca-2M.textquestion-answering1M<n<10M8 likes87 downloads2y agoHugging Face21uralstech /kcc-krishi-rag-sft-advisory-corpus KCC-Krishi RAG/SFT Advisory Corpus The KCC-Krishi RAG/SFT Advisory Corpus is a translated, quality-controlled, routing-aware research corpus derived from Kisan Call Centre records from the Government of India open-data ecosystem. It was created for: agricultural RAG research; supervised fine-tuning research; evidence-grounded response generation; safety-routing experiments; offline farmer-assistant prototyping; reproducible dataset and model-training experiments.… See the full description on the dataset page: https://huggingface.co/datasets/uralstech/kcc-krishi-rag-sft-advisory-corpus.texttext-generation100K<n<1M0 likes82 downloads2mo agoHugging Face22Azzindani /ID_REG_MD_RAG 📑 Indonesian Regulation Markdown RAG Dataset (ID_REG_MD_RAG) This repository contains a highly structured, Markdown-optimized collection of Indonesian Regulations (Peraturan Perundang-undangan). This dataset is specifically engineered to solve the "structure loss" problem often encountered when building Retrieval-Augmented Generation (RAG) systems for complex legal documents. 🏛️ 💡 The Concept: Structural Integrity for RAG Legal documents in Indonesia follow a… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/ID_REG_MD_RAG.tabulartext-generation100K<n<1M3 likes81 downloads7mo agoHugging Face23rgcmainhub /rage-core-1 Rage Core 1 Supervised fine-tuning dataset for Rage Core Gen 1 — developed by DevNameGelo, Powered by RGC Rage Gen Core. This dataset teaches the model to: understand user intent and ask clarifying questions when requirements are ambiguous, deliver complete production-oriented implementations, debug real-world problems, make precise code edits, operate as a coding agent through explicit tool calls, handle multilingual users, process multimodal inputs honestly, and refuse to… See the full description on the dataset page: https://huggingface.co/datasets/rgcmainhub/rage-core-1.texttext-generationn<1K0 likes81 downloads13d agoHugging Face24StarsMakeGalaxy /medical-device-regulatory-graft-rag-1350 🏥 Medical Device Regulatory & Clinical Compliance RAG Dataset (1,350 Samples) This dataset contains 1,350 highly curated, 100% LLM-synthesized RAG samples for training Small Language Models (SLMs: 1B–4B parameters) in high-stakes Medical Device Regulatory & Quality Compliance. Methodological Foundation: Pioneer / Prometheus Closed-Loop Curriculum Synthesis: Multi-slice curriculum covering 5 core operational failure modes. Elsevier Computer Standards & Interfaces… See the full description on the dataset page: https://huggingface.co/datasets/StarsMakeGalaxy/medical-device-regulatory-graft-rag-1350.texttext-generation1K<n<10K0 likes78 downloads20d agoHugging Face25flashserve /RAGPulse RAGPulse: A Real-World RAG Workload Trace to Optimize RAG Serving Systems 🌐 Github Link | 🤗 Workload Trace | 📑 Arxiv Paper | 🤖 How to use? RAGPulse is a real-world RAG workload trace collected from an university-wide Q&A service scenario. The system has been serving over 40,000 students and faculties since April 2024, providing intelligent policy Q&A services. The trace contains a total of 7,106 records entries, sampled from one week of our Q&A service. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/flashserve/RAGPulse.tabulartext-generation1K<n<10K2 likes75 downloads10mo agoHugging Face26f20180301 /loft-rag-hotpotqa-32k LOFT RAG - HotpotQA (32k) Dataset Description This dataset is part of the LOFT (Long-context Open Foundation Tasks) benchmark, specifically the RAG (Retrieval-Augmented Generation) task. Dataset: HotpotQA Context Length: 32k Task Type: RAG (Retrieval-Augmented Generation) Language: English Source: LOFT Benchmark (Google DeepMind) Dataset Structure Data Fields context (string): Full prompt context including corpus documents and few-shot… See the full description on the dataset page: https://huggingface.co/datasets/f20180301/loft-rag-hotpotqa-32k.textquestion-answeringn<1K0 likes75 downloads10mo agoHugging Face27code-rag-bench /mbppMBPP dataset annotated with ground-truth programming solutions, to enable evaluations for retrieval and retrieval-augmented code generation. Please refer to code-rag-bench for more details. texttext-generationn<1K1 likes73 downloads2y agoHugging Face28f20180301 /loft-rag-musique-32k LOFT RAG - MuSiQue (32k) Dataset Description This dataset is part of the LOFT (Long-context Open Foundation Tasks) benchmark, specifically the RAG (Retrieval-Augmented Generation) task. Dataset: MuSiQue Context Length: 32k Task Type: RAG (Retrieval-Augmented Generation) Language: English Source: LOFT Benchmark (Google DeepMind) Dataset Structure Data Fields context (string): Full prompt context including corpus documents and few-shot… See the full description on the dataset page: https://huggingface.co/datasets/f20180301/loft-rag-musique-32k.textquestion-answeringn<1K0 likes73 downloads10mo agoHugging Face29HiTZ /elkarhizketak-RAG Dataset Card for ElkarHizketak RAG and its Disruptor Variants Base and disruptor variants of ElkarHizketak, built to stress-test conversational RAG systems in Basque under realistic interaction patterns (conversational openings, topic shifts). Dataset Details Dataset Description This dataset extends ElkarHizketak with a base variant (rewritten opening queries, retrieval-needed labels, retrieved chunks) and disruptor variants that inject… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/elkarhizketak-RAG.tabularquestion-answering1K<n<10K1 likes73 downloads3mo agoHugging Face30oddadmix /arabic-rag-chat-8k-eval arabic-rag-chat-8k-eval Per-row evaluation artifacts for the 8,192-token Arabic multi-turn RAG models: the test split, every model's raw replies, every judge verdict, and the rendered report for each. Thirteen judged models, all scored on the same 1,651 prompts by the same judge at temperature 0.0, so the comparison below is like-for-like and can be recomputed offline without a GPU or a judge server. This is the measurement half of oddadmix/100M-8192-Nawah-dsv4; the training… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-8k-eval.tabularquestion-answeringn<1K0 likes72 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.