datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cold-cases
Collaborative Open Legal Data (COLD) - Cases
COLD Cases is a dataset of 8.3 million United States legal decisions with text and metadata, formatted as compressed parquet files. If you'd like to view a sample of the dataset formatted as JSON Lines, you can view one here
This dataset exists to support the open legal movement exemplified by projects like
Pile of Law and
LegalBench.
A key input to legal understanding projects is caselaw -- the published, precedential decisions of… See the full description on the dataset page: https://huggingface.co/datasets/harvard-lil/cold-cases.closure-challenge-v2-cfd-cases
Closure Challenge v2 — CFD case files (heavy)
Companion dataset to anon-closure-challenge-v2/closure-challenge-v2.
This repository hosts the full OpenFOAM cases (mesh, fields, run scripts), VTK volume snapshots, and DNS-native partitioned VTUs for the 14 test cases of the Closure Challenge v2 benchmark, together with the standardized training and validation sets under data/train/ and data/validation/.
The lightweight integral-profile portion needed for reproducing the scoring… See the full description on the dataset page: https://huggingface.co/datasets/anon-closure-challenge-v2/closure-challenge-v2-cfd-cases.Motius-Leaderboard-Cases
Motius Leaderboard Case Assets
This dataset stores compact browser assets for the all-case comparison pages in
the public Motius leaderboards. It is a
visualization companion, not a training or evaluation dataset.
Folder
Population
Comparison
m2t-humanml3d-smpl/
4,400
HumanML3D input motion with TM2T, MotionGPT, MotionGPT3, and VerMo captions
t2m-humanml3d-smpl/
4,042
HumanML3D selected captions with every released T2M output
babel-sequential-smpl/
1,295
BABEL GT… See the full description on the dataset page: https://huggingface.co/datasets/ZeyuLing/Motius-Leaderboard-Cases.RefWave-Cluster-RunsDutch-Rechtspraak-court-casesProstate-Anatomical-Edge-Cases
Prostate-Anatomical-Edge-Cases
Stress-Testing Pelvic Autosegmentation Algorithms Using Anatomical Edge Cases —
a TCIA collection of pelvic radiotherapy planning CT with manually contoured
organs at risk, curated so that most cases contain anatomy known to break
autosegmentation algorithms (Kanwar et al., Phys Imaging Radiat Oncol 2023).
Read before using — the name is misleading in two ways:
This is CT, not MRI. Despite "Prostate" in the name it is not a prostate
mpMRI/zonal… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/Prostate-Anatomical-Edge-Cases.TreeOil_Painting_ScientificJourney_Thailand_CaseStudy🧪 Tree Oil Painting: A Scientific Journey – Thailand Case Study
This dataset documents a rare and detailed forensic investigation of a mysterious 19th-century oil painting, known as The Tree Oil Painting, using scientific methods and AI-assisted analysis. Compiled in Thailand between 2015 and 2025, this work represents a grassroots effort to validate the painting’s origins through physical evidence, pigment mapping, synchrotron spectroscopy, and historical comparison.
🧩 Overview
Title: Tree… See the full description on the dataset page: https://huggingface.co/datasets/HaruthaiAi/TreeOil_Painting_ScientificJourney_Thailand_CaseStudy.india-tb-missed-cases-analysis
India TB Missed Cases Analysis & Living Model (2025)
🌟 Project Overview
This repository hosts a comprehensive, multi-method analytical framework designed to estimate and understand the "missing" millions of Tuberculosis (TB) cases in India. By integrating Bayesian statistics, Dimensionality Reduction (PCA), and Causal Inference (DAG), this project provides a high-resolution view of TB detection determinants across Indian states.
Core Analytical Pillars:… See the full description on the dataset page: https://huggingface.co/datasets/hssling/india-tb-missed-cases-analysis.ai-election-manipulation-cases
AI, Elections and Agency Transfer Evidence Index
Version 0.4.4 · released 21 August 2026 · research cutoff 12 August 2026
The dataset contains 6 documented-manipulation records, not 1,087 cases. Read the counts in this order:
1,087 relational rows -> 64 catalogue entries -> 10 core records
-> 8 incident-eligible records
-> 6 documented-manipulation records
The other two incident-eligible records are transparent contested-use… See the full description on the dataset page: https://huggingface.co/datasets/apol/ai-election-manipulation-cases.us-court-cases
Dataset Card for "us-court-cases"
More Information needed
guertin-mcro-forensic-corpus-federal-cases
Guertin MCRO Forensic Corpus: Federal Cases
Contents: 4 federal cases — 24-2662/ 24-2662, Matthew Guertin v. Hennepin County (Court of Appeals for the Eighth Circuit; 14 PDFs); 24-cv-02646/ 0:24-cv-02646, Guertin v. Hennepin County (District Court, D. Minnesota; 106 PDFs); 25-2476/ 25-2476, Matthew Guertin v. Tim Walz (Court of Appeals for the Eighth Circuit; 128 PDFs); 25-cv-02670/ 0:25-cv-02670, Guertin v. Walz (District Court, D. Minnesota; 136 PDFs).
Layout: one folder per… See the full description on the dataset page: https://huggingface.co/datasets/Matt1up/guertin-mcro-forensic-corpus-federal-cases.xiaoluo-openp2p-input-cases
Input-only gameplay cases
Up to 100 independently sampled 15-second cases for each game with readable input candidates. The fixed seed balances source recordings and samples across time bins. Previous visual-review results are not used in sample selection.
Visually unreviewed: subtitles, HUD complexity and visible interaction outcomes have not been screened. Key names describe recorded controls, not inferred game actions. Horizon Zero Dawn currently has no readable input… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/xiaoluo-openp2p-input-cases.asia-who-treatment-success-rate-hiv-positive-tb-cases
Treatment success rate: HIV-positive TB cases | Asia (WHO GHO)
🌏 551 observations · 45 Asia countries · 1999–2023 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 551 observations of Treatment success rate: HIV-positive TB cases data across 45 Asia countries, spanning 1999–2023, covering 1 distinct indicators.
About the source
Source: WHO Global Health Observatory
Publisher: World Health Organization
License: cc-by-4.0
Topic:… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-who-treatment-success-rate-hiv-positive-tb-cases.game-interaction-cases-1k-media
Game interaction gallery media
Original natural 15-second H.264 gameplay clips for Play / Field Notes.
The gallery contains 1000 labeled samples across GTA V, RDR2, Watch Dogs 2 and Hogwarts Legacy. Five similarity pairs involving seven samples remain under review; the gallery marks them explicitly. Original source metadata, event labels, and dataset revisions are in the paired Space's data/cases directory. Media repositories are split only when the Hub enforces a per-repository… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/game-interaction-cases-1k-media.ecthr_casesThe ECtHR Cases dataset is designed for experimentation of neural judgment prediction and rationale extraction considering ECtHR cases.cold-cases
Collaborative Open Legal Data (COLD) - Cases
COLD Cases is a dataset of 8.3 million United States legal decisions with text and metadata, formatted as compressed parquet files. If you'd like to view a sample of the dataset formatted as JSON Lines, you can view one here
This dataset exists to support the open legal movement exemplified by projects like
Pile of Law and
LegalBench.
A key input to legal understanding projects is caselaw -- the published, precedential decisions of… See the full description on the dataset page: https://huggingface.co/datasets/5-31-2024/cold-cases.safety-calibration-cases
Safety Calibration Cases
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and reproducibility… See the full description on the dataset page: https://huggingface.co/datasets/rmems/safety-calibration-cases.M3Retrieve_CaseStudyRetrievalasylum-casesmeasurement-axioms-cases
Measurement Axioms Cases
Source pinning
Frozen, source-pinned publication — not a live mirror of the canonical repository's main.
Source snapshot commit
350bb4cba4e5bc2d760db080aae52352a7041331 (Measurement Axioms v1.0.0)
Canonical current repository
https://github.com/halvrenofviryel/measurement-axioms
Export/publication date
2026-09-13 (first Hub commit of this repository)
Update policy
Counts are derived from this snapshot: 45 active… See the full description on the dataset page: https://huggingface.co/datasets/phionyx/measurement-axioms-cases.africa-who-treatment-success-rate-hiv-positive-tb-cases
Africa — WHO GHO: Treatment success rate: HIV-positive TB cases | Africa (World Health Organization)
Size category: n<1K - Formats: parquet - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-who-treatment-success-rate-hiv-positive-tb-cases.CaseSumm
Dataset Summary
This dataset was introduced in CaseSumm: A Large-Scale Dataset for Long-Context Summarization from U.S. Supreme Court Opinions.
The CaseSumm dataset consists of U.S. Supreme Court cases and their official summaries, called syllabuses, from the period 1815-2019. Syllabuses are written by an attorney employed by the Court and approved by the Justices. The syllabus is therefore the gold standard for summarizing majority opinions, and ideal for evaluating other summaries… See the full description on the dataset page: https://huggingface.co/datasets/ChicagoHAI/CaseSumm.ioi-test-cases
IOI
The International Olympiad in Informatics (IOI) is one of five international science olympiads (if you are familiar with AIME, IOI is the programming equivalent of IMO, for which the very best students who take part in AIME are invited) and tests a very select group of high school students (4 per country) in complex algorithmic problems.
The problems are extremely challenging, and the full test sets are available and released under a permissive (CC-BY) license. This means that… See the full description on the dataset page: https://huggingface.co/datasets/open-r1/ioi-test-cases.court_cases_azerbaijani
Court Cases Of The Republic Of Azerbaijan
This dataset consists of court cases from the Republic of Azerbaijan.
Overview
It was formed based on 1,200,000 court cases.
The data has been preliminarily normalized and split into sentences.
The dataset consists of 37 million sentences and approximately 500-600 million tokens.
Dataset Structure
Each row represents a single sentence extracted from a court case document.
Column
Type
Description
case_id… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/court_cases_azerbaijani.Dutch-Judiciary-Court-Cases-Netherlands-RechtspraakSupreme-Court-Cases-1830-2019
US Supreme Court Legal Corpus (1830–2019)
Overview
A comprehensive, production-ready AI training dataset containing 456,589 documents from 122,930 US Supreme Court cases spanning 190 years (1830–2019).
This corpus captures the full adversarial record — petitions for certiorari, respondent briefs, reply briefs, amicus curiae filings, appendices, oral argument transcripts, and opinions. It is one of the most complete collections of Supreme Court procedural and… See the full description on the dataset page: https://huggingface.co/datasets/OwnedByDanes/Supreme-Court-Cases-1830-2019.real_clinical_cases_of_Famous_Old_TCM_Doctors
TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors 数据集简介
TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors是一个包含了当代著名老中医临床病例的数据集。这些病例数据来源于《当代名老中医典型医案集》(Contemporary Famous Old Chinese Medicine Doctors' Typical Cases Collection)一书。该数据集收录了多位德高望重的老中医大家的真实门诊病历,涵盖了多种常见病和疑难杂症。每个病例都包括病情描述、辨证论治思路、具体治疗方药等宝贵的一手临床资料。这些医案凝聚了老一辈名医的智慧和经验,对于中医的传承发展和临床应用研究,都有重要价值。通过对这些案例的挖掘分析,能够总结老中医诊疗思维、理法方药的特点,为现代中医临床实践提供有益借鉴。
Introduction to TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors… See the full description on the dataset page: https://huggingface.co/datasets/TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors.airep-evidence-cases
AIREP Evidence Cases
Source pinning
Frozen, source-pinned publication — not a live mirror of the canonical repository's main.
Source snapshot commit
8a6c01ecce457aa94330c0ed7219e4c56ebfe771 (v0.2.0-beta.1) · frozen v0.1.2 at 44387bd43cc06ba656eaa7ff670be5c8e3220aca · publication-source review ff5c3551052251726c0ed878dcc23a44e305bd93
Canonical current repository
https://github.com/halvrenofviryel/ai-runtime-evidence-protocol
Export/publication date… See the full description on the dataset page: https://huggingface.co/datasets/phionyx/airep-evidence-cases.oalc_cases
Dataset Card for OALC Cases
This dataset contains the subset of court cases from the Open Australian Legal Corpus compiled by Umar Butler. Ben Chen regularly updates this dataset.
usaco-incorrect-test-cases
USACO Incorrect Test Cases
This dataset contains subtly incorrect outputs for USACO competitive programming problems, designed for reward hacking detection experiments.
📊 Dataset Summary
Total problems: 56 platinum-level USACO problems
File structure: Each problem has a directory {problem_id}/2_incorrect.out
Purpose: Detect when models special-case or hard-code test outputs
Generation method: Created using GPT-4.1 to produce subtly incorrect but plausible outputs… See the full description on the dataset page: https://huggingface.co/datasets/rmcc11/usaco-incorrect-test-cases.
