datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FinQA-hallucination-detection
FinQA Hallucination Detection
Dataset Summary
This dataset was created from a subset of the original FinQA dataset. For each user query (financial questions), we prompted an LLM to generate a response to this query based on provided context (financial statements and tables from the original FinQA).
Each generated LLM response is labeled based on whether it is correct or not. This dataset is thus useful for benchmarking reference-free LLM Eval and Hallucination… See the full description on the dataset page: https://huggingface.co/datasets/Cleanlab/FinQA-hallucination-detection.Phantom_Hallucination_Detection
Phantom: A Benchmark for Hallucination Detection in Financial Long-Context QA
Authors: Lanlan Ji, Dominic Seyler, Gunkirat Kaur, Manjunath Hegde, Koustuv Dasgupta, Bing Xiang
This is the repository containing the dataset for the submission mentioned above.
This dataset is designed for hallucination detection in language models. It includes multiple variants of the Phantom dataset with different token lengths (seed, 2k, 5K, 10K, 20K, 30K) for long context experiments , segments… See the full description on the dataset page: https://huggingface.co/datasets/seyled/Phantom_Hallucination_Detection.LLM-Hallucination-Detection-complex-mathematics
AIME Hallucination Detection Dataset
This dataset is created for detecting hallucinations in Large Language Models (LLMs), particularly focusing on complex mathematical problems. It can be used for tasks like model evaluation, fine-tuning, and research.
Dataset Details
Name: AIME Hallucination Detection Dataset
Format: CSV
Size: (add size, e.g., 10MB)
Files Included:
AIME-hallucination-detection-dataset.csv: Contains the dataset.
Content Description… See the full description on the dataset page: https://huggingface.co/datasets/tourist800/LLM-Hallucination-Detection-complex-mathematics.FinQA-hallucination-detection
FinQA Hallucination Detection
Dataset Summary
This dataset was created from a subset of the original FinQA dataset. For each user query (financial questions), we prompted an LLM to generate a response to this query based on provided context (financial statements and tables from the original FinQA).
Each generated LLM response is labeled based on whether it is correct or not. This dataset is thus useful for benchmarking reference-free LLM Eval and Hallucination… See the full description on the dataset page: https://huggingface.co/datasets/kankshith123/FinQA-hallucination-detection.Phantom_Hallucination_Detection
Phantom: A Benchmark for Hallucination Detection in Financial Long-Context QA
Authors: Lanlan Ji, Dominic Seyler, Gunkirat Kaur, Manjunath Hegde, Koustuv Dasgupta, Bing Xiang
This is the repository containing the dataset for the submission mentioned above.
This dataset is designed for hallucination detection in language models. It includes multiple variants of the Phantom dataset with different token lengths (seed, 2k, 5K, 10K, 20K, 30K) for long context experiments , segments… See the full description on the dataset page: https://huggingface.co/datasets/Anindita1979/Phantom_Hallucination_Detection.hallucination-detection
Dataset Summary
Hallucination Detection dataset is a specialized dataset designed to evaluate language models' tendency to hallucinate (generate factually incorrect or unsupported information) in the Earth Observation (EO) domain. Unlike typical QA datasets that focus on correctness, this dataset contains deliberately hallucinated answers with detailed annotations marking which portions of the text are hallucinated.
This dataset was introduced as part of the paper EVE: A… See the full description on the dataset page: https://huggingface.co/datasets/eve-esa/hallucination-detection.med-hallucination-detection
Medical hallucination detection
A dataset for training a small model to detect hallucinations in medical answers
and explain why, by checking each answer against the context it should be
grounded in. Each row is a (question, answer, context) triple with a row_type:
not_hallucinated -- the answer is grounded in its context.
hallucinated -- the answer is not (sourced separately; see below).
The not_hallucinated split (this build)
Derived from MedQuAD, a collection… See the full description on the dataset page: https://huggingface.co/datasets/Certops/med-hallucination-detection.AIME_Hallucination_Detection
AIME Hallucination Detection Dataset
This dataset is created for detecting hallucinations in Large Language Models (LLMs), particularly focusing on complex mathematical problems. It can be used for tasks like model evaluation, fine-tuning, and research.
Dataset Details
Name: AIME Hallucination Detection Dataset
Format: CSV
Size: (14.6 MB)
Files Included:
AIME-hallucination-detection-dataset.csv: Contains the dataset.
Content Description
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/tourist800/AIME_Hallucination_Detection.med-hallucination-detection-unfiltered
Medical hallucination detection (unfiltered)
A dataset for training a small model to detect hallucinations in medical answers
and explain why, by checking each answer against the context it should be
grounded in. Each row is a (question, answer, context) triple labelled row_type.
This is the unfiltered union of two sources: 7,464 grounded positives and
10,000 planted-hallucination negatives. It is the raw pool before sampling and
judging -- the downstream step samples from here… See the full description on the dataset page: https://huggingface.co/datasets/Certops/med-hallucination-detection-unfiltered.Hallucination-Detectioncot-oracle-hallucination-detectiontoolace-hallucination-span-detection
ToolACE Hallucination Span Detection
This dataset was generated for span-level hallucination detection in tool-calling dialogues.
It is derived from Team-ACE/ToolACE.
Main file:
generated_data/toolace_ragtruth_all_with_negatives_and_splits.jsonl
Each example contains a query, tool context, final answer, hallucination labels,
hallucination type, split, and tool metadata. Labels are character-level spans
in the answer.
Construction
We extract completed ToolACE… See the full description on the dataset page: https://huggingface.co/datasets/katsubakirill/toolace-hallucination-span-detection.llm-hallucination-detectionhallucination_detection_transformers
ToolACE Hallucination Dataset
This dataset was generated for the assignment Hallucination Detection in Tool Calling.
It is based on ToolACE tool-calling dialogues and uses a RAGTruth-style schema:
query: user question
context: tool output / grounding evidence
output: assistant final answer
hallucination_labels: character-level hallucination spans
Files:
File
Rows
Description
toolace_clean_ragtruth.jsonl
1347
clean ToolACE tool-use answers… See the full description on the dataset page: https://huggingface.co/datasets/HASSANI8046/hallucination_detection_transformers.
