datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RAGTruth-processed
RAGTruth Dataset
Dataset Description
Dataset Summary
The RAGTruth dataset is designed for evaluating hallucinations in text generation models, particularly in retrieval-augmented generation (RAG) contexts. It contains examples of model outputs along with expert annotations indicating whether the outputs contain hallucinations.
Dataset Structure
Each example contains:
A query/question
Context passages
Model output
Hallucination labels (evident… See the full description on the dataset page: https://huggingface.co/datasets/wandb/RAGTruth-processed.ragtruth-translated-hallucinations
RAGTruth Translated Hallucinations
Multilingual machine translation of
RAGTruth into 31 European languages,
preserving RAGTruth's word-level hallucination-span annotations. RAGTruth is a corpus of
LLM responses to retrieval-augmented generation (RAG) tasks in which humans marked the
exact spans that are hallucinated (unsupported by, or contradicting, the provided
context). Here both the RAG prompt and the response are translated into each target
language, and the annotated… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/ragtruth-translated-hallucinations.RAGTruth_test
RAGTruth test set
Dataset
Test split of RAGTruth dataset by ParticleMedia available from https://github.com/ParticleMedia/RAGTruth/tree/main/dataset
The dataset was published in RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models
Preprocessing
We kept only the test split of the original dataset
Joined response and source info files
Created the response level hallucination labels as described in the paper using binary… See the full description on the dataset page: https://huggingface.co/datasets/flowaicom/RAGTruth_test.toolace-tool-calling-hallucination-ragtruth
ToolACE-derived Tool-Calling Hallucination Dataset
This dataset was created for the course assignment Hallucination Detection in Tool Calling.
It is synthetic by design: starting from ToolACE-style tool-calling dialogues, we automatically inject three required hallucination types:
tool_contradiction
overgeneration
missing_tool
Each example follows a RAGTruth-like format:
query: user query
context: tool output
output: final model answer
hallucination_labels: span-level… See the full description on the dataset page: https://huggingface.co/datasets/marrita/toolace-tool-calling-hallucination-ragtruth.ragtruthRAGTruth-TR
RAGTruth-TR
newmindai/RAGTruth-TR is a Turkish-translated version of the wandb/RAGTruth-processed dataset.
It is designed for evaluating Retrieval-Augmented Generation (RAG) systems in Turkish, enabling research in hallucination detection, fact-checking, and response quality assessment.
Dataset Summary
Source Dataset: wandb/RAGTruth-processed
Target Language: Turkish
Purpose: Hallucination detection and RAG evaluation in Turkish NLP systems
License: MIT (inherits from… See the full description on the dataset page: https://huggingface.co/datasets/newmindai/RAGTruth-TR.eval_ragtruth-qa_SFT_gemma-4-E4B-it_S130104_epo3_89de_gens_T0_wfs0_s12345_mt512_nosftragtruth-qa_sftragtruth-qa_rm_organicragtruth-qa_perlragtruth-qa_processedragtruth-qa_final_test_seteval_ragtruth-qa_PERL_gemma-4-E4B-it_S130104_ace3_gens_T0_wfs0_s12345_mt512_sft1c9bb9ragtruth-qa_autoraterragtruth-qa_rm_synthetic_structeval_ragtruth-qa_gemma-4-E4B-it_gens_T0_wfs0_s12345_mt512_nosftragtruth-uncragtruth-de-translatedThe dataset is created from the RAGTruth dataset by translating it to German. We've used Gemma 3 27B for the translation.
The translation was done on a single A100 machine using VLLM as a server.
ragtruth-cn-translatedThe dataset is created from the RAGTruth dataset by translating it to Chinese. We've used Gemma 3 27B for the translation.
The translation was done on a single A100 machine using VLLM as a server.
ragtruth-pl-translatedThe dataset is created from the RAGTruth dataset by translating it to Polish. We've used Gemma 3 27B for the translation.
The translation was done on a single A100 machine using VLLM as a server.
RAGTruthPrefixesragtruth-fr-translatedThe dataset is created from the RAGTruth dataset by translating it to French. We've used Gemma 3 27B for the translation.
The translation was done on a single A100 machine using VLLM as a server.
ragtruth-hu-translatedThe dataset is created from the RAGTruth dataset by translating it to Hungarian. We've used Gemma 3 27B for the translation.
The translation was done on a single A100 machine using VLLM as a server.
ragtruth-es-translatedThe dataset is created from the RAGTruth dataset by translating it to Spanish. We've used Gemma 3 27B for the translation.
The translation was done on a single A100 machine using VLLM as a server.
ragtruth-it-translatedThe dataset is created from the RAGTruth dataset by translating it to Italian. We've used Gemma 3 27B for the translation.
The translation was done on a single A100 machine using VLLM as a server.
ragtruth_rm_organicragtruth-de-translated-manual-300ragtruth_sftragtruth_rm_synthetic_llmrag_truth_hallucination_binary
