datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swiss-caselaw
Swiss Case Law Dataset
1,050,000+ published decision records (~909,000 unique decisions) from Swiss federal, cantonal, and regulatory bodies.
Figures as of 2026-07-24 — refreshed daily; live counts at opencaselaw.ch.
Full text, structured metadata, extracted case-citation references, and daily updates. The dataset contains German, French, and Italian decisions; the export schema also reserves rm for Romansh.
Dataset Summary
The largest open collection of… See the full description on the dataset page: https://huggingface.co/datasets/voilaj/swiss-caselaw.Caselaw_Access_Project_embeddingsThis is an embeddings dataset for the Caselaw Access Project, created by a user named Endomorphosis.
Each caselaw entry is hashed with IPFS / multiformats, so retrieval of the document can be made over the IPFS / filecoin network
The ipfs content id "cid" is the primary key that links the dataset to the embeddings, should you want to retrieve from the dataset instead.
The dataset has been had embeddings generated with three models: thenlper/gte-small, Alibaba-NLP/gte-large-en-v1.5, and… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/Caselaw_Access_Project_embeddings.caselaw_access_project
Caselaw Access Project
Description
This dataset contains 6.7 million cases from the Caselaw Access Project and Court Listener.
The Caselaw Access Project consists of nearly 40 million pages of U.S. federal and state court decisions and judges’ opinions from the last 365 years.
In addition, Court Listener adds over 900 thousand cases scraped from 479 courts.
The Caselaw Access Project and Court Listener source legal data from a wide variety of resources such as the… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/caselaw_access_project.cold-cases
Collaborative Open Legal Data (COLD) - Cases
COLD Cases is a dataset of 8.3 million United States legal decisions with text and metadata, formatted as compressed parquet files. If you'd like to view a sample of the dataset formatted as JSON Lines, you can view one here
This dataset exists to support the open legal movement exemplified by projects like
Pile of Law and
LegalBench.
A key input to legal understanding projects is caselaw -- the published, precedential decisions of… See the full description on the dataset page: https://huggingface.co/datasets/harvard-lil/cold-cases.Caselaw_Access_Project_embeddingsOriginal Repository:
https://huggingface.co/datasets/justicedao/Caselaw_Access_Project_embeddings/
This is an embeddings dataset for the Caselaw Access Project, created by a user named Endomorphosis.
Each caselaw entry is hashed with IPFS / multiformats, so retrieval of the document can be made over the IPFS / filecoin network
The ipfs content id "cid" is the primary key that links the dataset to the embeddings, should you want to retrieve from the dataset instead.
The dataset has been had… See the full description on the dataset page: https://huggingface.co/datasets/laion/Caselaw_Access_Project_embeddings.Caselaw-Access-Projectcanadian-case-law
A2AJ Canadian Case Law
Last updated: 2026-09-20
Maintainer: Access to Algorithmic Justice (A2AJ)
Dataset Summary
The A2AJ Canadian Case Law dataset provides bulk, open-access full-text decisions from Canadian courts and tribunals.
Each row corresponds to a single case and contains the English and French versions of the decision where both are publicly available.
Rows also include citation network fields (cases_cited, cases_citing, and citing_cases_count)… See the full description on the dataset page: https://huggingface.co/datasets/a2aj/canadian-case-law.case-law
The Case-law, centralizing legal decisions for better use, a community Dataset.
The Case-law Dataset is a comprehensive collection of legal decisons from various countries, centralized in a common format. This dataset aims to improve the development of legal AI models by providing a standardized, easily accessible corpus of global legal documents.
Join us in our mission to make AI more accessible and understandable for the legal world, ensuring that the power of language models… See the full description on the dataset page: https://huggingface.co/datasets/HFforLegal/case-law.nsw-caselaw-qa-embedcaseholdCaseHOLD (Case Holdings On Legal Decisions) is a law dataset comprised of over 53,000+ multiple choice questions to identify the relevant holding of a cited case.nsw-caselaw-chunkedAILA_casedocs
AILACasedocs
An MTEB dataset
Massive Text Embedding Benchmark
The task is to retrieve the case document that most closely matches or is most relevant to the scenario described in the provided query.
Task category
t2t
Domains
Legal, Written
Reference
https://zenodo.org/records/4063986
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["AILACasedocs"])
evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/AILA_casedocs.indian-case-laws
Indian Case Laws
Open Indian case-law data for AI, search, and legal research.
This dataset is part of the KanoonGPT Open Legal Data Initiative - an effort to make Indian legal data easier to access, trace, and build on for open-source research, legal tech, and production AI systems.
KanoonGPT is building structured Indian legal data and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.
Repository: KanoonGPT/indian-case-laws… See the full description on the dataset page: https://huggingface.co/datasets/KanoonGPT/indian-case-laws.caselaw_access_project_filtered
Caselaw Access Project
Description
This dataset contains 6.7 million cases from the Caselaw Access Project and Court Listener.
The Caselaw Access Project consists of nearly 40 million pages of U.S. federal and state court decisions and judges’ opinions from the last 365 years.
In addition, Court Listener adds over 900 thousand cases scraped from 479 courts.
The Caselaw Access Project and Court Listener source legal data from a wide variety of resources such as the Harvard… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/caselaw_access_project_filtered.RefWave-Cluster-Runsnsw-caselaw-qaus-caselaw-oh
Ohio Case Law
Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them.
Source & credit — Free Law Project / CourtListener
Every opinion in this dataset comes from the Free Law Project / CourtListener bulk export of
2026-06-30 (10,798,347 opinions). CourtListener… See the full description on the dataset page: https://huggingface.co/datasets/docketx/us-caselaw-oh.Dutch-Rechtspraak-court-casesProstate-Anatomical-Edge-Cases
Prostate-Anatomical-Edge-Cases
Stress-Testing Pelvic Autosegmentation Algorithms Using Anatomical Edge Cases —
a TCIA collection of pelvic radiotherapy planning CT with manually contoured
organs at risk, curated so that most cases contain anatomy known to break
autosegmentation algorithms (Kanwar et al., Phys Imaging Radiat Oncol 2023).
Read before using — the name is misleading in two ways:
This is CT, not MRI. Despite "Prostate" in the name it is not a prostate
mpMRI/zonal… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/Prostate-Anatomical-Edge-Cases.evalarc-casebook
EvalArc Casebook
The same 93.75% score can pass one acceptance gate and fail another.
Inspect the rules, actual failed checks and original Docker records in a
filterable table. This is the data companion to the
interactive evidence lab.
In the default suite_jobs view, compare support-partial and
support-protected. Both use the same frozen defective policy, score 93.75%
and fully resolve 0/2 attempts. The deliberately permissive rule accepts partial
progress; the rule requiring… See the full description on the dataset page: https://huggingface.co/datasets/glayguo/evalarc-casebook.ai-election-manipulation-cases
AI, Elections and Agency Transfer Evidence Index
Version 0.4.4 · released 21 August 2026 · research cutoff 12 August 2026
The dataset contains 6 documented-manipulation records, not 1,087 cases. Read the counts in this order:
1,087 relational rows -> 64 catalogue entries -> 10 core records
-> 8 incident-eligible records
-> 6 documented-manipulation records
The other two incident-eligible records are transparent contested-use… See the full description on the dataset page: https://huggingface.co/datasets/apol/ai-election-manipulation-cases.us-caselaw
US Caselaw
Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them.
The American case-law record in one dataset — published opinions from the Free Law Project /
CourtListener bulk export of 2026-06-30: full text, all opinion types (majority, concurring,
dissenting, per curiam… See the full description on the dataset page: https://huggingface.co/datasets/docketx/us-caselaw.us-court-cases
Dataset Card for "us-court-cases"
More Information needed
us-caselaw-wi
Wisconsin Case Law
Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them.
Full text of 72,350 Wisconsin appellate opinion documents from the public record,
sliced from the Free Law Project / CourtListener bulk export of 2026-06-30.
Court coverage (2 court ids, explicit allowlist —… See the full description on the dataset page: https://huggingface.co/datasets/docketx/us-caselaw-wi.ua-case-outcome-6m
Ukrainian Court Decisions: Case Outcome Prediction (6.7M)
The largest publicly available dataset of Ukrainian court decisions for case outcome prediction, extracted from the State Court Decisions Registry (EDRSR). Contains 6,690,284 substantive decisions from civil and commercial courts spanning 2008--2026, with temporal splits across three wartime epochs.
Overview
Ukraine's EDRSR is one of the world's largest open judicial databases, containing 100M+ judicial… See the full description on the dataset page: https://huggingface.co/datasets/overthelex/ua-case-outcome-6m.Venus_Case_Tempus-caselaw-ca
California Case Law
Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them.
Full text of 242,989 California appellate opinion documents from the public record. The base
slice is 242,831 documents from the Free Law Project / CourtListener bulk export of 2026-06-30.… See the full description on the dataset page: https://huggingface.co/datasets/docketx/us-caselaw-ca.Seal-Tools
Seal-Tools
This Huggingface repository contains the dataset generated in Seal-Tools: Self-Instruct Tool Learning Dataset for Agent Tuning and Detailed Benchmark.
Abstract
Seal-Tools contains self-instruct API-like tools. Seal-Tools not only offers a large
number of tools, but also includes instances
which demonstrate the practical application
of tools. Seeking to generate data on a large
scale while ensuring reliability, we propose a
self-instruct method to generate… See the full description on the dataset page: https://huggingface.co/datasets/casey-martin/Seal-Tools.texas-caselaw
Texas Case Law
Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them.
Source & credit — Free Law Project / CourtListener
Every opinion in this dataset comes from the Free Law Project / CourtListener bulk export of
2026-06-30 (10,798,347 opinions). CourtListener… See the full description on the dataset page: https://huggingface.co/datasets/docketx/texas-caselaw.guertin-mcro-forensic-corpus-federal-cases
Guertin MCRO Forensic Corpus: Federal Cases
Contents: 4 federal cases — 24-2662/ 24-2662, Matthew Guertin v. Hennepin County (Court of Appeals for the Eighth Circuit; 14 PDFs); 24-cv-02646/ 0:24-cv-02646, Guertin v. Hennepin County (District Court, D. Minnesota; 106 PDFs); 25-2476/ 25-2476, Matthew Guertin v. Tim Walz (Court of Appeals for the Eighth Circuit; 128 PDFs); 25-cv-02670/ 0:25-cv-02670, Guertin v. Walz (District Court, D. Minnesota; 136 PDFs).
Layout: one folder per… See the full description on the dataset page: https://huggingface.co/datasets/Matt1up/guertin-mcro-forensic-corpus-federal-cases.
