CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Cleanlab /FinQA-hallucination-detection FinQA Hallucination Detection Dataset Summary This dataset was created from a subset of the original FinQA dataset. For each user query (financial questions), we prompted an LLM to generate a response to this query based on provided context (financial statements and tables from the original FinQA). Each generated LLM response is labeled based on whether it is correct or not. This dataset is thus useful for benchmarking reference-free LLM Eval and Hallucination… See the full description on the dataset page: https://huggingface.co/datasets/Cleanlab/FinQA-hallucination-detection.text1K<n<10K3 likes1.4k downloads2y agoHugging Face02seyled /Phantom_Hallucination_Detection Phantom: A Benchmark for Hallucination Detection in Financial Long-Context QA Authors: Lanlan Ji, Dominic Seyler, Gunkirat Kaur, Manjunath Hegde, Koustuv Dasgupta, Bing Xiang This is the repository containing the dataset for the submission mentioned above. This dataset is designed for hallucination detection in language models. It includes multiple variants of the Phantom dataset with different token lengths (seed, 2k, 5K, 10K, 20K, 30K) for long context experiments , segments… See the full description on the dataset page: https://huggingface.co/datasets/seyled/Phantom_Hallucination_Detection.text10K<n<100K2 likes247 downloads11mo agoHugging Face03tourist800 /LLM-Hallucination-Detection-complex-mathematics AIME Hallucination Detection Dataset This dataset is created for detecting hallucinations in Large Language Models (LLMs), particularly focusing on complex mathematical problems. It can be used for tasks like model evaluation, fine-tuning, and research. Dataset Details Name: AIME Hallucination Detection Dataset Format: CSV Size: (add size, e.g., 10MB) Files Included: AIME-hallucination-detection-dataset.csv: Contains the dataset. Content Description… See the full description on the dataset page: https://huggingface.co/datasets/tourist800/LLM-Hallucination-Detection-complex-mathematics.tabularn<1K1 likes73 downloads2y agoHugging Face04kankshith123 /FinQA-hallucination-detection FinQA Hallucination Detection Dataset Summary This dataset was created from a subset of the original FinQA dataset. For each user query (financial questions), we prompted an LLM to generate a response to this query based on provided context (financial statements and tables from the original FinQA). Each generated LLM response is labeled based on whether it is correct or not. This dataset is thus useful for benchmarking reference-free LLM Eval and Hallucination… See the full description on the dataset page: https://huggingface.co/datasets/kankshith123/FinQA-hallucination-detection.text1K<n<10K0 likes60 downloads4mo agoHugging Face05Anindita1979 /Phantom_Hallucination_Detection Phantom: A Benchmark for Hallucination Detection in Financial Long-Context QA Authors: Lanlan Ji, Dominic Seyler, Gunkirat Kaur, Manjunath Hegde, Koustuv Dasgupta, Bing Xiang This is the repository containing the dataset for the submission mentioned above. This dataset is designed for hallucination detection in language models. It includes multiple variants of the Phantom dataset with different token lengths (seed, 2k, 5K, 10K, 20K, 30K) for long context experiments , segments… See the full description on the dataset page: https://huggingface.co/datasets/Anindita1979/Phantom_Hallucination_Detection.text10K<n<100K1 likes53 downloads5mo agoHugging Face06eve-esa /hallucination-detection Dataset Summary Hallucination Detection dataset is a specialized dataset designed to evaluate language models' tendency to hallucinate (generate factually incorrect or unsupported information) in the Earth Observation (EO) domain. Unlike typical QA datasets that focus on correctness, this dataset contains deliberately hallucinated answers with detailed annotations marking which portions of the text are hallucinated. This dataset was introduced as part of the paper EVE: A… See the full description on the dataset page: https://huggingface.co/datasets/eve-esa/hallucination-detection.texttext-classification1K<n<10K0 likes46 downloads5mo agoHugging Face07Certops /med-hallucination-detection Medical hallucination detection A dataset for training a small model to detect hallucinations in medical answers and explain why, by checking each answer against the context it should be grounded in. Each row is a (question, answer, context) triple with a row_type: not_hallucinated -- the answer is grounded in its context. hallucinated -- the answer is not (sourced separately; see below). The not_hallucinated split (this build) Derived from MedQuAD, a collection… See the full description on the dataset page: https://huggingface.co/datasets/Certops/med-hallucination-detection.textquestion-answering1K<n<10K0 likes37 downloads1mo agoHugging Face08tourist800 /AIME_Hallucination_Detection AIME Hallucination Detection Dataset This dataset is created for detecting hallucinations in Large Language Models (LLMs), particularly focusing on complex mathematical problems. It can be used for tasks like model evaluation, fine-tuning, and research. Dataset Details Name: AIME Hallucination Detection Dataset Format: CSV Size: (14.6 MB) Files Included: AIME-hallucination-detection-dataset.csv: Contains the dataset. Content Description The dataset… See the full description on the dataset page: https://huggingface.co/datasets/tourist800/AIME_Hallucination_Detection.tabularn<1K0 likes26 downloads2y agoHugging Face09Certops /med-hallucination-detection-unfiltered Medical hallucination detection (unfiltered) A dataset for training a small model to detect hallucinations in medical answers and explain why, by checking each answer against the context it should be grounded in. Each row is a (question, answer, context) triple labelled row_type. This is the unfiltered union of two sources: 7,464 grounded positives and 10,000 planted-hallucination negatives. It is the raw pool before sampling and judging -- the downstream step samples from here… See the full description on the dataset page: https://huggingface.co/datasets/Certops/med-hallucination-detection-unfiltered.textquestion-answering10K<n<100K0 likes26 downloads1mo agoHugging Face10Karamjit99 /Hallucination-Detection0 likes24 downloads1y agoHugging Face11ceselder /cot-oracle-hallucination-detection0 likes14 downloads7mo agoHugging Face12katsubakirill /toolace-hallucination-span-detection ToolACE Hallucination Span Detection This dataset was generated for span-level hallucination detection in tool-calling dialogues. It is derived from Team-ACE/ToolACE. Main file: generated_data/toolace_ragtruth_all_with_negatives_and_splits.jsonl Each example contains a query, tool context, final answer, hallucination labels, hallucination type, split, and tool metadata. Labels are character-level spans in the answer. Construction We extract completed ToolACE… See the full description on the dataset page: https://huggingface.co/datasets/katsubakirill/toolace-hallucination-span-detection.token-classification0 likes8 downloads4mo agoHugging Face13nuriasane /llm-hallucination-detectiontextn<1K0 likes7 downloads5mo agoHugging Face14HASSANI8046 /hallucination_detection_transformers ToolACE Hallucination Dataset This dataset was generated for the assignment Hallucination Detection in Tool Calling. It is based on ToolACE tool-calling dialogues and uses a RAGTruth-style schema: query: user question context: tool output / grounding evidence output: assistant final answer hallucination_labels: character-level hallucination spans Files: File Rows Description toolace_clean_ragtruth.jsonl 1347 clean ToolACE tool-use answers… See the full description on the dataset page: https://huggingface.co/datasets/HASSANI8046/hallucination_detection_transformers.tabular1K<n<10K1 likes6 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.