CoolFace
Datasetpublic

docketx/us-caselaw-ny

New York Case Law Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them. Full text of 963,045 New York appellate opinion documents from the public record, sliced from the Free Law Project / CourtListener bulk export of 2026-06-30. Court coverage (5 court ids, explicit allowlist —… See the full description on the dataset page: https://huggingface.co/datasets/docketx/us-caselaw-ny.

sourceHugging Facecc0-1.0updated 3d agoView on Hugging Face
0likes236downloads
Dataset Card

New York Case Law

Code & tools: github.com/docketxlegal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them.

Full text of 963,045 New York appellate opinion documents from the public record, sliced from the Free Law Project / CourtListener bulk export of 2026-06-30.

Court coverage (5 court ids, explicit allowlist — never prefix-matched): ny, nyappdiv, nyappterm, nychanct, nyoytermct.

Lead opinions and separate opinions; documents under 1 characters excluded. Slice sha256 906d6641a1452ff6836255bd1dc0855bb00964d4ba7082cf9d68f6cb868d99e1 (also in slice.manifest.json).

Composition

Measured on this exact slice: 407,798 documents (42.3%) carry 2,000+ characters of text, 344,057 (35.7%) carry 200-1,999, and 211,190 (21.9%) are under 200 characters. The short ones are real records, not errors: certiorari denials, one-line orders, and judgment entries. They are kept deliberately (min_chars = 1) so this is the complete public record rather than a filtered subset. Filter on len(text) if you want only substantive opinions.

Format

Rows follow docketx record v1 (one opinion per line): id, doc_type, jurisdiction, title, text, source, license, retrieved_at, plus citation, court, date, and extra (CourtListener opinion/cluster ids, opinion type). The same format is used across every docketx dataset.

Provenance and license

Judicial opinions are edicts of government: uncopyrightable works of the public domain (Banks v. Manchester, 128 U.S. 244 (1888); Georgia v. Public.Resource.Org, 590 U.S. 255 (2020)). This packaging is released under CC0 1.0. Source: CourtListener bulk export (2026-06-30), Free Law Project — https://free.law. This dataset redistributes public-domain court text; it adds no annotation and asserts no rights.

Load it

python
from datasets import load_dataset
ds = load_dataset("docketx/us-caselaw-ny")

Part of the DocketRouter legal corpora: https://huggingface.co/docketx