datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gpt-failure-cases-dataset
Dataset Summary
This dataset contains a curated collection of medical question–answer pairs designed to evaluate large language models (LLMs) such as GPT-4 and GPT-5 on their ability to provide factually correct responses. The dataset highlights failure cases (hallucinations) where both models struggled, making it a valuable benchmark for studying factual consistency and reliability in AI-generated medical content.
Each entry consists of:
question: A natural language medical query.… See the full description on the dataset page: https://huggingface.co/datasets/ehe07/gpt-failure-cases-dataset.LLM-Failure-Cases
Codatta LLM Failure Cases (Expert Critiques)
Overview
Codatta LLM Failure Cases is a specialized adversarial dataset designed to highlight and analyze scenarios where state-of-the-art Large Language Models (LLMs) produce incorrect, hallucinatory, or logically flawed responses.
This dataset originates from Codatta's "Airdrop Season 1" campaign, a crowdsourced data intelligence initiative where participants were tasked with finding prompts that caused leading LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Humanbased-AI/LLM-Failure-Cases.
