datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Supreme-Court-Cases-1830-2019
US Supreme Court Legal Corpus (1830–2019)
Overview
A comprehensive, production-ready AI training dataset containing 456,589 documents from 122,930 US Supreme Court cases spanning 190 years (1830–2019).
This corpus captures the full adversarial record — petitions for certiorari, respondent briefs, reply briefs, amicus curiae filings, appendices, oral argument transcripts, and opinions. It is one of the most complete collections of Supreme Court procedural and… See the full description on the dataset page: https://huggingface.co/datasets/OwnedByDanes/Supreme-Court-Cases-1830-2019.ode-enterprise-use-cases
ODE Enterprise Use Case Dataset
15,000 labeled enterprise use cases spanning 31 modules, 215 submodules, 8 industry verticals, 5 channels, and 12 business personas.
Published by Llewellyn Systems Inc — builders of ODE, the Operating System for Decision & Enterprise.
Attribution Required
This dataset is licensed under CC-BY-4.0. You are free to use, share, and adapt this dataset for any purpose — including commercial — as long as you give appropriate credit.
How… See the full description on the dataset page: https://huggingface.co/datasets/LlewellynSystems/ode-enterprise-use-cases.gpt-failure-cases-dataset
Dataset Summary
This dataset contains a curated collection of medical question–answer pairs designed to evaluate large language models (LLMs) such as GPT-4 and GPT-5 on their ability to provide factually correct responses. The dataset highlights failure cases (hallucinations) where both models struggled, making it a valuable benchmark for studying factual consistency and reliability in AI-generated medical content.
Each entry consists of:
question: A natural language medical query.… See the full description on the dataset page: https://huggingface.co/datasets/ehe07/gpt-failure-cases-dataset.ode-enterprise-use-cases
ODE Enterprise Use Case Dataset
15,000 labeled enterprise use cases spanning 31 modules, 215 submodules, 8 industry verticals, 5 channels, and 12 business personas.
Published by Llewellyn Systems Inc — builders of ODE, the Operating System for Decision & Enterprise.
Attribution Required
This dataset is licensed under CC-BY-4.0. You are free to use, share, and adapt this dataset for any purpose — including commercial — as long as you give appropriate credit.
How… See the full description on the dataset page: https://huggingface.co/datasets/LlewellynSystemsInc/ode-enterprise-use-cases.runpod_multi_model_think_content_casestudiesLLM-Failure-Cases
Codatta LLM Failure Cases (Expert Critiques)
Overview
Codatta LLM Failure Cases is a specialized adversarial dataset designed to highlight and analyze scenarios where state-of-the-art Large Language Models (LLMs) produce incorrect, hallucinatory, or logically flawed responses.
This dataset originates from Codatta's "Airdrop Season 1" campaign, a crowdsourced data intelligence initiative where participants were tasked with finding prompts that caused leading LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Humanbased-AI/LLM-Failure-Cases.packetcourt-golden-cases
PacketCourt Golden Cases
A small evidence-first evaluation set for auditing front-of-pack claims against
the text printed on the same Indian packaged-food label.
Each record contains:
front-label claim text
back-label evidence text
expected claims and conservative verdicts
expected persuasion-gap concepts
expected deterministic date or whole-packet calculations
The initial set is intentionally small and hand-audited. It is a regression and
demonstration asset, not a… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/packetcourt-golden-cases.
