failure-cases
gpt-failure-cases-dataset
Dataset Summary
This dataset contains a curated collection of medical question–answer pairs designed to evaluate large language models (LLMs) such as GPT-4 and GPT-5 on their ability to provide factually correct responses. The dataset highlights failure cases (hallucinations) where both models struggled, making it a valuable benchmark for studying factual consistency and reliability in AI-generated medical content.
Each entry consists of:
question: A natural language medical query.… See the full description on the dataset page: https://huggingface.co/datasets/ehe07/gpt-failure-cases-dataset.libero_ds_failure_cases_addedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 98,
"total_frames": 13097,
"total_tasks": 10,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:98"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/WillMandil001/libero_ds_failure_cases_added.qwen3-vl-failure-cases
Qwen3-VL-2B-Instruct Failure Analysis Dataset
📊 Dataset Overview
This dataset contains 10 diverse failure cases identified while testing the Qwen3-VL-2B-Instruct vision-language model. Each example captures a specific type of error, providing valuable insights for targeted fine-tuning.
Failure Category
Count
Examples
Time Reading
2
Clock misreading (11:55 vs 10:10; 3:35 vs 10:35)
Counting
2
Remote buttons (3 vs 0); Strawberries (4 vs 1)
Negation… See the full description on the dataset page: https://huggingface.co/datasets/TasneemSelim/qwen3-vl-failure-cases.LLM-Failure-Cases
Codatta LLM Failure Cases (Expert Critiques)
Overview
Codatta LLM Failure Cases is a specialized adversarial dataset designed to highlight and analyze scenarios where state-of-the-art Large Language Models (LLMs) produce incorrect, hallucinatory, or logically flawed responses.
This dataset originates from Codatta's "Airdrop Season 1" campaign, a crowdsourced data intelligence initiative where participants were tasked with finding prompts that caused leading LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Humanbased-AI/LLM-Failure-Cases.FF-Nanbeige4-3B-Base-failure-cases
I- MODEL USED
Published in December 2025, the selected model is Nanbeige4-3B-Base accessible at: https://huggingface.co/Nanbeige/Nanbeige4-3B-Base.
It is a 3-billion-parameter (which is within the required 0.6B–6B range) foundation model (base model) from the fourth generation of the Nanbeige LLM series.
It demonstrates that a compact architecture can deliver strong performance when paired with rigorous improvements to data quality and training methods.
II- MODEL EXPLORATION AND BLIND SPOTS… See the full description on the dataset page: https://huggingface.co/datasets/Choukouriyah/FF-Nanbeige4-3B-Base-failure-cases.africa-who-retreatment-cases-treatment-after-failure
Africa — WHO GHO: Retreatment cases: treatment after failure (pulmonary smear and/or culture positive) | Africa (World Health Organization)
Size category: n<1K - Formats: parquet - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-who-retreatment-cases-treatment-after-failure.
