datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swiss-caselaw
Swiss Case Law Dataset
1,050,000+ published decision records (~909,000 unique decisions) from Swiss federal, cantonal, and regulatory bodies.
Figures as of 2026-07-24 — refreshed daily; live counts at opencaselaw.ch.
Full text, structured metadata, extracted case-citation references, and daily updates. The dataset contains German, French, and Italian decisions; the export schema also reserves rm for Romansh.
Dataset Summary
The largest open collection of… See the full description on the dataset page: https://huggingface.co/datasets/voilaj/swiss-caselaw.Caselaw_Access_Project_embeddingsThis is an embeddings dataset for the Caselaw Access Project, created by a user named Endomorphosis.
Each caselaw entry is hashed with IPFS / multiformats, so retrieval of the document can be made over the IPFS / filecoin network
The ipfs content id "cid" is the primary key that links the dataset to the embeddings, should you want to retrieve from the dataset instead.
The dataset has been had embeddings generated with three models: thenlper/gte-small, Alibaba-NLP/gte-large-en-v1.5, and… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/Caselaw_Access_Project_embeddings.swiss-caselaw-web-ui
Swiss Case Law Open Dataset
962,724 published decisions from Swiss federal, cantonal, and regulatory bodies.
Full text, structured metadata, and daily updates. The March 20, 2026 snapshot contains German, French, and Italian decisions; the export schema also reserves rm for Romansh.
What this is
A structured, searchable archive of Swiss court decisions — from the Federal Supreme Court (BGer) down to cantonal courts in all 26 cantons. Every decision includes the full… See the full description on the dataset page: https://huggingface.co/datasets/ArneH/swiss-caselaw-web-ui.Caselaw_Access_Project_JSON
The Caselaw Access Project
In collaboration with Ravel Law, Harvard Law Library digitized over 40 million U.S. court decisions consisting of 6.7 million cases from the last 360 years into a dataset that is widely accessible to use. Access a bulk download of the data through the Caselaw Access Project API (CAPAPI): https://case.law/caselaw/
Find more information about accessing state and federal written court decisions of common law through the bulk data service documentation here:… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/Caselaw_Access_Project_JSON.caselaw_access_project
Caselaw Access Project
Description
This dataset contains 6.7 million cases from the Caselaw Access Project and Court Listener.
The Caselaw Access Project consists of nearly 40 million pages of U.S. federal and state court decisions and judges’ opinions from the last 365 years.
In addition, Court Listener adds over 900 thousand cases scraped from 479 courts.
The Caselaw Access Project and Court Listener source legal data from a wide variety of resources such as the… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/caselaw_access_project.cold-cases
Collaborative Open Legal Data (COLD) - Cases
COLD Cases is a dataset of 8.3 million United States legal decisions with text and metadata, formatted as compressed parquet files. If you'd like to view a sample of the dataset formatted as JSON Lines, you can view one here
This dataset exists to support the open legal movement exemplified by projects like
Pile of Law and
LegalBench.
A key input to legal understanding projects is caselaw -- the published, precedential decisions of… See the full description on the dataset page: https://huggingface.co/datasets/harvard-lil/cold-cases.Caselaw_Access_Project_embeddingsOriginal Repository:
https://huggingface.co/datasets/justicedao/Caselaw_Access_Project_embeddings/
This is an embeddings dataset for the Caselaw Access Project, created by a user named Endomorphosis.
Each caselaw entry is hashed with IPFS / multiformats, so retrieval of the document can be made over the IPFS / filecoin network
The ipfs content id "cid" is the primary key that links the dataset to the embeddings, should you want to retrieve from the dataset instead.
The dataset has been had… See the full description on the dataset page: https://huggingface.co/datasets/laion/Caselaw_Access_Project_embeddings.Caselaw-Access-Projectcanadian-case-law
A2AJ Canadian Case Law
Last updated: 2026-09-20
Maintainer: Access to Algorithmic Justice (A2AJ)
Dataset Summary
The A2AJ Canadian Case Law dataset provides bulk, open-access full-text decisions from Canadian courts and tribunals.
Each row corresponds to a single case and contains the English and French versions of the decision where both are publicly available.
Rows also include citation network fields (cases_cited, cases_citing, and citing_cases_count)… See the full description on the dataset page: https://huggingface.co/datasets/a2aj/canadian-case-law.case-law
The Case-law, centralizing legal decisions for better use, a community Dataset.
The Case-law Dataset is a comprehensive collection of legal decisons from various countries, centralized in a common format. This dataset aims to improve the development of legal AI models by providing a standardized, easily accessible corpus of global legal documents.
Join us in our mission to make AI more accessible and understandable for the legal world, ensuring that the power of language models… See the full description on the dataset page: https://huggingface.co/datasets/HFforLegal/case-law.closure-challenge-v2-cfd-cases
Closure Challenge v2 — CFD case files (heavy)
Companion dataset to anon-closure-challenge-v2/closure-challenge-v2.
This repository hosts the full OpenFOAM cases (mesh, fields, run scripts), VTK volume snapshots, and DNS-native partitioned VTUs for the 14 test cases of the Closure Challenge v2 benchmark, together with the standardized training and validation sets under data/train/ and data/validation/.
The lightweight integral-profile portion needed for reproducing the scoring… See the full description on the dataset page: https://huggingface.co/datasets/anon-closure-challenge-v2/closure-challenge-v2-cfd-cases.nsw-caselaw-qa-embedMotius-Leaderboard-Cases
Motius Leaderboard Case Assets
This dataset stores compact browser assets for the all-case comparison pages in
the public Motius leaderboards. It is a
visualization companion, not a training or evaluation dataset.
Folder
Population
Comparison
m2t-humanml3d-smpl/
4,400
HumanML3D input motion with TM2T, MotionGPT, MotionGPT3, and VerMo captions
t2m-humanml3d-smpl/
4,042
HumanML3D selected captions with every released T2M output
babel-sequential-smpl/
1,295
BABEL GT… See the full description on the dataset page: https://huggingface.co/datasets/ZeyuLing/Motius-Leaderboard-Cases.caseholdCaseHOLD (Case Holdings On Legal Decisions) is a law dataset comprised of over 53,000+ multiple choice questions to identify the relevant holding of a cited case.nsw-caselaw-chunkedopf_small_case2000_goc
Data download through hfApi
Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times.
opf_small_case57_ieee
Data download through hfApi
Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times.
opf_small_case10000_gocopf_small_case14_ieee
Data download through hfApi
Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times.
opfdata_case2000_gocopf_small_case30_ieee
Data download through hfApi
Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times.
AILA_casedocs
AILACasedocs
An MTEB dataset
Massive Text Embedding Benchmark
The task is to retrieve the case document that most closely matches or is most relevant to the scenario described in the provided query.
Task category
t2t
Domains
Legal, Written
Reference
https://zenodo.org/records/4063986
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["AILACasedocs"])
evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/AILA_casedocs.pf_small_case10000_gocopfdata_case30_ieeepf_small_case2000_goc
Data download through hfApi
Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times.
opf_small_case500_goc
Data download through hfApi
Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times.
indian-case-laws
Indian Case Laws
Open Indian case-law data for AI, search, and legal research.
This dataset is part of the KanoonGPT Open Legal Data Initiative - an effort to make Indian legal data easier to access, trace, and build on for open-source research, legal tech, and production AI systems.
KanoonGPT is building structured Indian legal data and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.
Repository: KanoonGPT/indian-case-laws… See the full description on the dataset page: https://huggingface.co/datasets/KanoonGPT/indian-case-laws.opf_small_case118_ieee
Data download through hfApi
Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times.
caselaw_access_project_filtered
Caselaw Access Project
Description
This dataset contains 6.7 million cases from the Caselaw Access Project and Court Listener.
The Caselaw Access Project consists of nearly 40 million pages of U.S. federal and state court decisions and judges’ opinions from the last 365 years.
In addition, Court Listener adds over 900 thousand cases scraped from 479 courts.
The Caselaw Access Project and Court Listener source legal data from a wide variety of resources such as the Harvard… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/caselaw_access_project_filtered.opfdata_case118_ieee
