CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OwnedByDanes /Supreme-Court-Cases-1830-2019 US Supreme Court Legal Corpus (1830–2019) Overview A comprehensive, production-ready AI training dataset containing 456,589 documents from 122,930 US Supreme Court cases spanning 190 years (1830–2019). This corpus captures the full adversarial record — petitions for certiorari, respondent briefs, reply briefs, amicus curiae filings, appendices, oral argument transcripts, and opinions. It is one of the most complete collections of Supreme Court procedural and… See the full description on the dataset page: https://huggingface.co/datasets/OwnedByDanes/Supreme-Court-Cases-1830-2019.tabulartext-generation10K<n<100K0 likes131 downloads5mo agoHugging Face02LlewellynSystems /ode-enterprise-use-cases ODE Enterprise Use Case Dataset 15,000 labeled enterprise use cases spanning 31 modules, 215 submodules, 8 industry verticals, 5 channels, and 12 business personas. Published by Llewellyn Systems Inc — builders of ODE, the Operating System for Decision & Enterprise. Attribution Required This dataset is licensed under CC-BY-4.0. You are free to use, share, and adapt this dataset for any purpose — including commercial — as long as you give appropriate credit. How… See the full description on the dataset page: https://huggingface.co/datasets/LlewellynSystems/ode-enterprise-use-cases.texttext-classification10K<n<100K0 likes58 downloads7mo agoHugging Face03ehe07 /gpt-failure-cases-dataset Dataset Summary This dataset contains a curated collection of medical question–answer pairs designed to evaluate large language models (LLMs) such as GPT-4 and GPT-5 on their ability to provide factually correct responses. The dataset highlights failure cases (hallucinations) where both models struggled, making it a valuable benchmark for studying factual consistency and reliability in AI-generated medical content. Each entry consists of: question: A natural language medical query.… See the full description on the dataset page: https://huggingface.co/datasets/ehe07/gpt-failure-cases-dataset.textquestion-answering1K<n<10K0 likes42 downloads1y agoHugging Face04LlewellynSystemsInc /ode-enterprise-use-cases ODE Enterprise Use Case Dataset 15,000 labeled enterprise use cases spanning 31 modules, 215 submodules, 8 industry verticals, 5 channels, and 12 business personas. Published by Llewellyn Systems Inc — builders of ODE, the Operating System for Decision & Enterprise. Attribution Required This dataset is licensed under CC-BY-4.0. You are free to use, share, and adapt this dataset for any purpose — including commercial — as long as you give appropriate credit. How… See the full description on the dataset page: https://huggingface.co/datasets/LlewellynSystemsInc/ode-enterprise-use-cases.texttext-classification10K<n<100K0 likes34 downloads7mo agoHugging Face05DataTonic /runpod_multi_model_think_content_casestudiestexttext-generation1K<n<10K1 likes30 downloads2y agoHugging Face06Humanbased-AI /LLM-Failure-Cases Codatta LLM Failure Cases (Expert Critiques) Overview Codatta LLM Failure Cases is a specialized adversarial dataset designed to highlight and analyze scenarios where state-of-the-art Large Language Models (LLMs) produce incorrect, hallucinatory, or logically flawed responses. This dataset originates from Codatta's "Airdrop Season 1" campaign, a crowdsourced data intelligence initiative where participants were tasked with finding prompts that caused leading LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Humanbased-AI/LLM-Failure-Cases.textquestion-answeringn<1K1 likes24 downloads10mo agoHugging Face07Coltuc2026 /Romanian-Legal-Cases-2026 Romanian Legal Cases Dataset - 2026 (Litigii Bancare & Comerciale) Acest dataset conține mii de spețe anonimizate din practica juridică a Cabinetului Avocat Marius Vicențiu Coltuc, specializat în litigii bancare și drept comercial în România. Descriere Setul de date este destinat cercetării în domeniul LegalTech și antrenării modelelor de limbaj (LLM) pentru înțelegerea terminologiei juridice românești și a logicii judiciare curente. Fiecare intrare include:… See the full description on the dataset page: https://huggingface.co/datasets/Coltuc2026/Romanian-Legal-Cases-2026.text-classification1K<n<10K0 likes12 downloads4mo agoHugging Face08build-small-hackathon /packetcourt-golden-cases PacketCourt Golden Cases A small evidence-first evaluation set for auditing front-of-pack claims against the text printed on the same Indian packaged-food label. Each record contains: front-label claim text back-label evidence text expected claims and conservative verdicts expected persuasion-gap concepts expected deterministic date or whole-packet calculations The initial set is intentionally small and hand-audited. It is a regression and demonstration asset, not a… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/packetcourt-golden-cases.texttext-classificationn<1K0 likes12 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.