datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RAGTruth_test
RAGTruth test set
Dataset
Test split of RAGTruth dataset by ParticleMedia available from https://github.com/ParticleMedia/RAGTruth/tree/main/dataset
The dataset was published in RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models
Preprocessing
We kept only the test split of the original dataset
Joined response and source info files
Created the response level hallucination labels as described in the paper using binary… See the full description on the dataset page: https://huggingface.co/datasets/flowaicom/RAGTruth_test.ragtrutheval_ragtruth-qa_SFT_gemma-4-E4B-it_S130104_epo3_89de_gens_T0_wfs0_s12345_mt512_nosftragtruth-qa_sftragtruth-qa_rm_organicragtruth-qa_perlragtruth-qa_processedragtruth-qa_final_test_seteval_ragtruth-qa_PERL_gemma-4-E4B-it_S130104_ace3_gens_T0_wfs0_s12345_mt512_sft1c9bb9ragtruth-qa_autoraterragtruth-qa_rm_synthetic_structtoolace_ragtruth_synthetic
ToolACE RAGTruth style synthetic hallucination dataset
This repository contains a synthetic span annotated hallucination detection dataset for tool calling responses.
The data was derived from Team ACE ToolACE conversations. Only grounded instances with a user query, tool call, tool output, and final assistant answer were used.
Motivation
The dataset targets hallucination detection in tool augmented generation, where a final assistant response should be grounded in the… See the full description on the dataset page: https://huggingface.co/datasets/Ali-Bhai/toolace_ragtruth_synthetic.eval_ragtruth-qa_gemma-4-E4B-it_gens_T0_wfs0_s12345_mt512_nosftragtruth-unctoolace-ragtruth-style-hallucinations
ToolACE RAGTruth-style Tool Hallucination Dataset
This dataset was built from ToolACE tool-use dialogues and converted into a RAGTruth-style format for hallucination detection in tool calling.
Task
Given:
query: user query
context: tool response
output: final assistant answer
the goal is to classify whether the answer is grounded in the tool output or belongs to one of three hallucination types.
Labels
clean
tool_output_conflict
overgeneration… See the full description on the dataset page: https://huggingface.co/datasets/Ali-Bhai/toolace-ragtruth-style-hallucinations.ragtruth_rm_organicragtruth_sftragtruth_rm_synthetic_llmragtruth_processedragtruth_autorater
