datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
us-court-opinions-dockets-judges-dataset
US Court Opinions Metadata, Dockets & Judges (CourtListener)
10M opinion clusters, 70M dockets and 16K judges from official CourtListener / Free Law Project bulk data as metadata + derived-signals tables — citation graph, company litigation profiles; no opinion full text.
Part of the DataForge Open Data program — full production
packages, free for academic and personal use. Canonical dataset page:
https://data.zalize.com/datasets/us-court-opinions-dockets-judges-dataset… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/us-court-opinions-dockets-judges-dataset.kl3m-data-dockets
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dockets.kl3m-dockets-sampleus-dockets
US Dockets — the case-level record
Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them.
34,290,148 docket records from 2,289 U.S. courts: the case, the court, the
number, the dates, the judge. This is the index of American litigation — what was filed, where, when
and by whom it… See the full description on the dataset page: https://huggingface.co/datasets/docketx/us-dockets.fc-imm-dockets
Refugee Law Lab: Federal Court Immigration Dockets
Dataset Summary
The Refugee Law Lab supports bulk open-access to Canadian legal data to facilitate research and advocacy.
Bulk open-access helps avoid asymmetrical access-to-justice and amplification of marginalization that
results when commercial actors leverage proprietary
legal datasets for profit -- a particular concern in the border control setting.
This is the dataset is an updated version of a dataset used for a… See the full description on the dataset page: https://huggingface.co/datasets/refugee-law-lab/fc-imm-dockets.
