CoolFace
Datasetpublic

docketx/us-caselaw-oh

Ohio Case Law Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them. Source & credit — Free Law Project / CourtListener Every opinion in this dataset comes from the Free Law Project / CourtListener bulk export of 2026-06-30 (10,798,347 opinions). CourtListener… See the full description on the dataset page: https://huggingface.co/datasets/docketx/us-caselaw-oh.

sourceHugging Facecc0-1.0updated 3d agoView on Hugging Face
0likes290downloads
Dataset Card

Ohio Case Law

Code & tools: github.com/docketxlegal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them.

Source & credit — Free Law Project / CourtListener

Every opinion in this dataset comes from the Free Law Project / CourtListener bulk export of 2026-06-30 (10,798,347 opinions). CourtListener gathered, cleaned, deduplicated and published that record; Free Law Project — the 501(c)(3) non-profit behind it, building free public access to American law since 2010 — pays for it. We sliced and reformatted; they did the hard part. Go to the source: <https://www.courtlistener.com>.

Free Law Project is a non-profit, and it is worth supporting: <https://free.law/donate/>. If this data is useful to you, credit Free Law Project and CourtListener too, and consider going straight to the source: CourtListener runs its own API and MCP server.

Full text of 270,910 Ohio appellate opinion documents from the public record, sliced from the Free Law Project / CourtListener bulk export of 2026-06-30.

Court coverage (2 court ids, explicit allowlist — never prefix-matched): ohio, ohioctapp.

Lead opinions and separate opinions; documents under 1 characters excluded. Slice sha256 96bfaa98b0764334e5b02c1debaff0a2905e5d0640b03175d97acfad8662c3a8 (also in slice.manifest.json).

Composition

Measured on this exact slice: 200,553 documents (74.0%) carry 2,000+ characters of text, 39,430 (14.6%) carry 200-1,999, and 30,927 (11.4%) are under 200 characters. The short ones are real records, not errors: certiorari denials, one-line orders, and judgment entries. They are kept deliberately (min_chars = 1) so this is the complete public record rather than a filtered subset. Filter on len(text) if you want only substantive opinions.

Format

Rows follow docketx record v1 (one opinion per line): id, doc_type, jurisdiction, title, text, source, license, retrieved_at, plus citation, court, date, and extra (CourtListener opinion/cluster ids, opinion type). The same format is used across every docketx dataset.

Provenance and license

Judicial opinions are edicts of government: uncopyrightable works of the public domain (Banks v. Manchester, 128 U.S. 244 (1888); Georgia v. Public.Resource.Org, 590 U.S. 255 (2020)). This packaging is released under CC0 1.0. Source: CourtListener bulk export (2026-06-30), Free Law Project — https://free.law. This dataset redistributes public-domain court text; it adds no annotation and asserts no rights.

Load it

python
from datasets import load_dataset
ds = load_dataset("docketx/us-caselaw-oh")

Part of the DocketRouter legal corpora: https://huggingface.co/docketx