CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NovachronoAI /RAG-Grounded-QA-188k 🎯 RAG Grounded QA 186K The Anti-Hallucination Dataset Teach language models to answer from context — or shut up trying. Built by NovachronoAI — Precision AI for the real world. Full Dataset (186K) · 20K Subset · Schema · Sources · Usage Guide 🧠 Why This Dataset Exists Most QA datasets teach models what to say. This one also teaches them when to stay silent. RAG (Retrieval-Augmented Generation) systems have a fatal flaw: the model hallucinates when… See the full description on the dataset page: https://huggingface.co/datasets/NovachronoAI/RAG-Grounded-QA-188k.tabularquestion-answering100K<n<1M0 likes167 downloads7mo agoHugging Face02CJJones /Wikipedia_RAG_QA_Classification 🏛️ Wikipedia RAG QA Dataset for Retrieval-Augmented Generation Training 📊 Dataset Description This dataset contains 300,000+ validated model-generated responses to Wikipedia content, specifically designed for Retrieval-Augmented Generation (RAG) applications and SQL database insertion tasks. Generated by Jeeney AI Reloaded 207M GPT with specialized RAG tuning. 🖥️ Demo Interface: Discord Live Chat Demo on Discord: https://discord.gg/Xe9tHFCS9h The full CJ… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Wikipedia_RAG_QA_Classification.tabulartext-generation100K<n<1M1 likes112 downloads6mo agoHugging Face03Azzindani /ID_REG_MD_RAG 📑 Indonesian Regulation Markdown RAG Dataset (ID_REG_MD_RAG) This repository contains a highly structured, Markdown-optimized collection of Indonesian Regulations (Peraturan Perundang-undangan). This dataset is specifically engineered to solve the "structure loss" problem often encountered when building Retrieval-Augmented Generation (RAG) systems for complex legal documents. 🏛️ 💡 The Concept: Structural Integrity for RAG Legal documents in Indonesia follow a… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/ID_REG_MD_RAG.tabulartext-generation100K<n<1M3 likes81 downloads7mo agoHugging Face04flashserve /RAGPulse RAGPulse: A Real-World RAG Workload Trace to Optimize RAG Serving Systems 🌐 Github Link | 🤗 Workload Trace | 📑 Arxiv Paper | 🤖 How to use? RAGPulse is a real-world RAG workload trace collected from an university-wide Q&A service scenario. The system has been serving over 40,000 students and faculties since April 2024, providing intelligent policy Q&A services. The trace contains a total of 7,106 records entries, sampled from one week of our Q&A service. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/flashserve/RAGPulse.tabulartext-generation1K<n<10K2 likes75 downloads10mo agoHugging Face05HiTZ /elkarhizketak-RAG Dataset Card for ElkarHizketak RAG and its Disruptor Variants Base and disruptor variants of ElkarHizketak, built to stress-test conversational RAG systems in Basque under realistic interaction patterns (conversational openings, topic shifts). Dataset Details Dataset Description This dataset extends ElkarHizketak with a base variant (rewritten opening queries, retrieval-needed labels, retrieved chunks) and disruptor variants that inject… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/elkarhizketak-RAG.tabularquestion-answering1K<n<10K1 likes73 downloads3mo agoHugging Face06oddadmix /arabic-rag-chat-8k-eval arabic-rag-chat-8k-eval Per-row evaluation artifacts for the 8,192-token Arabic multi-turn RAG models: the test split, every model's raw replies, every judge verdict, and the rendered report for each. Thirteen judged models, all scored on the same 1,651 prompts by the same judge at temperature 0.0, so the comparison below is like-for-like and can be recomputed offline without a GPU or a judge server. This is the measurement half of oddadmix/100M-8192-Nawah-dsv4; the training… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-8k-eval.tabularquestion-answeringn<1K0 likes72 downloads1mo agoHugging Face07dev-jonathanb /cs50-educational-rag CS50 Pedagogical RAG Dataset 📜 Dataset Description This repository contains the data artifacts for the undergraduate thesis, which explores the use of a pedagogical chatbot with Retrieval-Augmented Generation (RAG) for Harvard's CS50: Introduction to Computer Science course. The project involved several stages of data processing, from raw content collection to the generation and curation of a high-quality evaluation dataset. To ensure full transparency and… See the full description on the dataset page: https://huggingface.co/datasets/dev-jonathanb/cs50-educational-rag.tabularquestion-answeringn<1K0 likes53 downloads1y agoHugging Face08GXMZU /llm-rag-agent-papers llm-rag-agent-papers Research papers on LLM, RAG, and AI Agents - Knowledge base for RAG pipeline Dataset Structure This dataset contains three subsets: llm: Large Language Model related content rag: Retrieval-Augmented Generation related content agent: AI Agent related content Usage from datasets import load_dataset # Load all subsets dataset = load_dataset("GXMZU/llm-rag-agent-papers") # Load specific subset llm_data =… See the full description on the dataset page: https://huggingface.co/datasets/GXMZU/llm-rag-agent-papers.tabulartext-generation1K<n<10K3 likes51 downloads9mo agoHugging Face09oddadmix /arabic-rag-support-25K Arabic RAG customer-support scenarios (27,927 rows) Synthetic Modern Standard Arabic customer-support scenarios for training small RAG answerers, distilled from unsloth/gemma-4-31B-it-NVFP4 on a local vLLM. Built as the training set for oddadmix/Nawah-50M-RAG-Support. Each row: a customer question + the knowledge-base chunks of one fictional company (products, prices, policies, FAQ entries) + the ideal grounded agent answer. One generation request invents one company KB and 4 QA… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-support-25K.tabularquestion-answering10K<n<100K0 likes49 downloads1mo agoHugging Face10rajistics /rag-qa-arena RAG QA Arena Annotated Dataset A comprehensive multi-domain question-answering dataset with citation annotations designed for evaluating Retrieval-Augmented Generation (RAG) systems, featuring faithful answers with proper source attribution across 6 specialized domains. 🎯 Dataset Overview This annotated version of the RAG QA Arena dataset includes citation information and gold document IDs, making it ideal for evaluating not just answer accuracy but also answer grounding… See the full description on the dataset page: https://huggingface.co/datasets/rajistics/rag-qa-arena.tabularquestion-answering10K<n<100K0 likes41 downloads1y agoHugging Face11proxectonos /dog-rag DOG-RAG: A Galician Benchmark for Legal Retrieval-Augmented Generation Click to expand Dataset description Dataset Structure Example Entry Categories Source Documents Dataset Versions Examples Additional information Acknowledgements Cite this dataset Dataset description This dataset contains question–answer triplets derived from publications of the Diario Oficial de Galicia (DOG), the official gazette of the autonomous community of Galicia (Spain).… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/dog-rag.tabularquestion-answeringn<1K0 likes39 downloads4mo agoHugging Face12oddadmix /arabic-rag-chat-grpo-5K Arabic multi-turn RAG conversations — GRPO pool (5,259 conversations) The reinforcement-learning half of oddadmix/arabic-rag-chat-30K: same generator, same validator, same schema, disjoint companies. It exists so GRPO explores fresh knowledge bases instead of taking a second pass over material the SFT already memorised. conversations turns companies this pool 5,259 14,018 309 Company-disjointness is exact and verified: this pool shares zero company_id values… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-grpo-5K.tabularquestion-answering1K<n<10K1 likes35 downloads1mo agoHugging Face13AGmind /agmind-rag-splitter-ru-data RU Context-Aware Document Split Датасет (teacher-distillation) для обучения русского context-aware сплиттера документов для RAG. Каждый пример учит модель где резать документ на самодостаточные смысловые чанки, держа таблицы и код целыми. Использован для модели AGmind/agmind-rag-splitter-ru. Код генерации и обучения: github.com/botAGI/AGmind-ML. Формат (Alpaca JSONL) { "instruction": "Раздели документ на смысловые части для системы поиска (RAG)...", "input":… See the full description on the dataset page: https://huggingface.co/datasets/AGmind/agmind-rag-splitter-ru-data.tabulartext-generation10K<n<100K0 likes34 downloads2mo agoHugging Face14DiscoPosse /RAGPulse RAGPulse: A Real-World RAG Workload Trace to Optimize RAG Serving Systems 🌐 Github Link | 🤗 Workload Trace | 📑 Arxiv Paper | 🤖 How to use? RAGPulse is a real-world RAG workload trace collected from an university-wide Q&A service scenario. The system has been serving over 40,000 students and faculties since April 2024, providing intelligent policy Q&A services. The trace contains a total of 7,106 records entries, sampled from one week of our Q&A service. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/DiscoPosse/RAGPulse.tabulartext-generation1K<n<10K0 likes30 downloads3mo agoHugging Face15omaressam1111 /multi-tafseer-quran-rag Quran Tafseer RAG Dataset A structured Arabic dataset of Quranic tafseer collected from eight classical and modern tafseer books.The dataset contains verse-aligned tafseer passages designed for Retrieval-Augmented Generation (RAG) systems and Arabic NLP research. Each record links a Quran verse with its corresponding tafseer explanation from one of the tafseer books and includes rich metadata such as surah information, tafseer source, and embedding-ready text. The dataset was… See the full description on the dataset page: https://huggingface.co/datasets/omaressam1111/multi-tafseer-quran-rag.tabularquestion-answering10K<n<100K7 likes26 downloads5mo agoHugging Face16DivyanshuSingh96 /aimi-anime-rag-dataset-sample 🎌 Ultimate Anime Dataset (8,248 Entries) | 1917-2025 A meticulously curated collection spanning 108 years of anime history Love this dataset and the Anime Receipts concept? You can download the complete project via the links below: 🚀 Unlock the Full Potential Product What You Get Get It Here Tier 1 8,248 Anime Dataset (Parquet) Tier 2 Full AiMi Recommendation System (Backend + UI) Tier 3 Ultimate AiMi Recommendation System + AiMi Anime… See the full description on the dataset page: https://huggingface.co/datasets/DivyanshuSingh96/aimi-anime-rag-dataset-sample.imagetext-retrievaln<1K5 likes24 downloads10mo agoHugging Face17Karmane /enterprise-rag-internal-knowledge-search-benchmark-sample Enterprise RAG and Internal Knowledge Search Benchmark Dataset -- Free Evaluation Sample This dataset packages a synthetic internal company workspace and an evidence-linked benchmark table into one product for teams building enterprise RAG systems, internal search assistants, knowledge-base copilots, and deep-search evaluation pipelines. The benchmark is designed around a realistic fictional company, AsteraOps Cloud, with multiple departments, renamed projects, stale roadmaps… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/enterprise-rag-internal-knowledge-search-benchmark-sample.tabulartext-generationn<1K2 likes23 downloads4mo agoHugging Face18oddadmix /arabic-rag-chat-30K Arabic multi-turn RAG customer-support conversations (31,294 conversations) Synthetic Modern Standard Arabic customer-support conversations for training small Arabic RAG assistants. The bulk was distilled from gemini-3.1-flash-lite via the Batch API; a first 2.4% came from unsloth/gemma-4-31B-it-NVFP4 on a local vLLM server before the run was moved off-GPU. Both teachers were given the same prompts and the same validator. Each row is one conversation of 1-5 rounds over one… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-30K.tabularquestion-answering10K<n<100K0 likes23 downloads1mo agoHugging Face19CJJones /Cosmopedia_QA_RAG_JSON_SQLiteThe full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com Want more? 🚀 Get the AI Startup Bundle from Gumroad. 🖥️ Demo Interface: Discord Discord: https://discord.gg/Xe9tHFCS9h **Custom RAG QA generation services can be made available for paying customers to process internal documentation. DM me on Discord if you are interested.Jeeney AI GPT Reloaded 207M/Cosmopedia Model Outputs Dataset Dataset Description This dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Cosmopedia_QA_RAG_JSON_SQLite.tabulartext-generation10K<n<100K2 likes15 downloads7mo agoHugging Face20JacktheLander /fungi-rag-agent-sft-database Fungi RAG Agent SFT Database This dataset contains the supervised fine-tuning database used to train the JacktheLander/smollm2-1.7b-fungi-rag-agent-distill-lora-gguf adapter. It was built for the project-owned fungi RAG learning system: JacktheLander/FunghiResearchAgent The examples teach a small local language model to follow the system's agent contract: use rag.search for evidence-backed mycology answers, use safety.review for wild-mushroom edibility, field-identification… See the full description on the dataset page: https://huggingface.co/datasets/JacktheLander/fungi-rag-agent-sft-database.tabulartext-generation1K<n<10K0 likes15 downloads4mo agoHugging Face21varun500 /adaptive_rag_hotpotqa Adaptive RAG HotpotQA Dataset This dataset is a processed version of HotpotQA designed for training Adaptive Retrieval-Augmented Generation (RAG) systems. Features input: The input text for the model output: The target output text retrieval_label: Whether retrieval is needed (0/1) hop: The reasoning hop number (1 or 2) type: The type of example (multi_hop_qa, single_hop_qa, multi_hop_gating, etc.) metadata: Additional information about the example including: answer:… See the full description on the dataset page: https://huggingface.co/datasets/varun500/adaptive_rag_hotpotqa.tabularquestion-answeringn<1K0 likes12 downloads1y agoHugging Face22Karmane /enterprise-rag-internal-knowledge-search-benchmarkgated Enterprise RAG and Internal Knowledge Search Benchmark Dataset This dataset packages a synthetic internal company workspace and an evidence-linked benchmark table into one product for teams building enterprise RAG systems, internal search assistants, knowledge-base copilots, and deep-search evaluation pipelines. The benchmark is designed around a realistic fictional company, AsteraOps Cloud, with multiple departments, renamed projects, stale roadmaps, support escalations… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/enterprise-rag-internal-knowledge-search-benchmark.tabulartext-generationn<1K0 likes9 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.