datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hallucination-reduction-dpo-100k
Hallucination Reduction DPO (100K)
100,000 DPO preference pairs training LLMs to stay within knowledge bounds. The chosen response is accurate and appropriately uncertain; the rejected response is confident but wrong — fabricated statistics, fake citations, wrong facts, overclaimed certainty.
Motivation
Hallucination is the #1 reliability concern blocking enterprise LLM adoption. Models fail in predictable patterns:
Inventing specific statistics with false… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/hallucination-reduction-dpo-100k.algozee_rag-based-hallucination-reduction-in-llms
RAG-Based Hallucination Reduction in LLMs
Introduction to Large Language Models and Hallucination Problem
Dataset Info
Source: Kaggle
Original Size: 0.17 MB
Kaggle Downloads: 43
Files: 1
Files
llm_rag_dataset_6k.csv.csv
Mirrored from Kaggle
