datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turkish-court-decisions
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/mrfg/turkish-court-decisions.us-court-opinions-dockets-judges-dataset
US Court Opinions Metadata, Dockets & Judges (CourtListener)
10M opinion clusters, 70M dockets and 16K judges from official CourtListener / Free Law Project bulk data as metadata + derived-signals tables — citation graph, company litigation profiles; no opinion full text.
Part of the DataForge Open Data program — full production
packages, free for academic and personal use. Canonical dataset page:
https://data.zalize.com/datasets/us-court-opinions-dockets-judges-dataset… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/us-court-opinions-dockets-judges-dataset.2026-08-14-courtroom
synth courtroom run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth courtroom run — per-stage snapshots (resumable generation cache)
date_generated
20260815_201700
constitution
constitutions/claude_distilled_09_principles_mid_20260804/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ b992089ffec3dbc23ba676ef5c1ebad319937daa
models
per-stage models — see manifest.json… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-14-courtroom.azerbaijan-court-data
Azerbaijan Court System Dataset
The most comprehensive open dataset of Azerbaijan's judicial system — 1.64 million structured records and 1.54 million court decision PDFs (~160 GB) covering court decisions, active cases, scheduled hearings, court registries, judges, lawyers, and mediator organizations.
Built for AI engineers, legal tech startups, and researchers who need real-world legal data at scale.
Quick Start
Load with Hugging Face datasets
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/ismatsamadov/azerbaijan-court-data.indian-court-decisions
Indian Court Decisions
A large-scale dataset of Indian court decisions with full text, metadata, and outcome labels covering the Supreme Court of India and 25 High Courts (1950–2026).
Dataset Summary
Config
Train
Validation
Test
Total
high_courts
11,682,776
1,459,319
1,457,934
14,600,029
supreme_court
40,044
4,990
5,019
50,053
Total
14,650,082
This is one of the largest publicly available legal NLP datasets, containing over 14.6 million… See the full description on the dataset page: https://huggingface.co/datasets/overthelex/indian-court-decisions.turkish-court-decisions
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/Alptekinege/turkish-court-decisions.russian-supreme-court-plenum-acts
Plenum Resolutions of the Supreme Court of Russia (1961–2026)
Every act published in the «Постановления Пленума» section of the Russian Supreme Court's
website: 1,504 records — 1,503 plenum resolutions plus 1 meeting
agenda — with full texts, metadata and the court's original attachments. Coverage
1961–2026; completeness verified against the court's own index at collection time
(the section reported exactly 1,504 documents).
Постановления Пленума ВС РФ — руководящие разъяснения… See the full description on the dataset page: https://huggingface.co/datasets/Alexey5676/russian-supreme-court-plenum-acts.pl-court-raw
Dataset Card for JuDDGES/pl-court-raw
Dataset Summary
The dataset consists of Polish Court judgments available at https://orzeczenia.ms.gov.pl/, containing full content of the judgments along with metadata sourced from official API and extracted from the judgment contents. This dataset contains raw data. For instruction dataset see JuDDGES/pl-court-instruct. For graph dataset see JuDDGES/pl-court-graph.
Supported Tasks and Leaderboards
The dataset can be… See the full description on the dataset page: https://huggingface.co/datasets/JuDDGES/pl-court-raw.saos-polish-court-judgments
SAOS Speeches Corpus — Orzeczenia Sądów Polskich
Korpus orzeczeń sądowych z Systemu Analizy Orzeczeń Sądowych (SAOS), wygenerowany z oficjalnego API SAOS (www.saos.org.pl/api/dump/judgments).
Statystyki
Metryka
Wartość
Sądy
sądy powszechne, Sąd Najwyższy, Naczelny Sąd Administracyjny, Trybunał Konstytucyjny
Typy orzeczeń
wyroki, postanowienia, uzasadnienia, zarządzenia, uchwały
Rekordy
296,692
Znaki
5,895,441,856
Słowa
855,425,796
Tokeny… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/saos-polish-court-judgments.turkish-court-decisions
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/Gyrevortex/turkish-court-decisions.turkish-court-decisions-duplicate
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/serdarsrts/turkish-court-decisions-duplicate.indian-court-decisions
Indian Court Decisions
A large-scale dataset of Indian court decisions with full text, metadata, and outcome labels covering the Supreme Court of India and 25 High Courts (1950–2026).
Dataset Summary
Config
Train
Validation
Test
Total
high_courts
11,682,776
1,459,319
1,457,934
14,600,029
supreme_court
40,044
4,990
5,019
50,053
Total
14,650,082
This is one of the largest publicly available legal NLP datasets, containing over 14.6 million… See the full description on the dataset page: https://huggingface.co/datasets/rtarun789/indian-court-decisions.sf_criminal_court
San Francisco Criminal Court Data
Linked criminal-court records for San Francisco County, combining scraped Superior Court docket data, District Attorney open-data feeds, and a charge-disposition spreadsheet that the Court produced only after sustained pressure under California Rules of Court rule 10.500.
Tables
File
Rows
Description
cases.parquet
77,406
Case-level records: case number, case ID, defendant name, filing date
register_of_actions.parquet
776,728… See the full description on the dataset page: https://huggingface.co/datasets/jamiequint/sf_criminal_court.pl-court-raw-enriched
Polish Court Judgments Raw (Enriched)
Polish court judgments enriched with Gemini-extracted factual_state and legal_state fields.
Dataset Description
This dataset is an enriched version of JuDDGES/pl-court-raw with additional fields extracted using Google Gemini 2.5 Pro.
New Fields
Core Extracted Fields
Field
Type
Description
factual_state
string
Objective narrative of facts (stan faktyczny) - the factual circumstances forming the basis… See the full description on the dataset page: https://huggingface.co/datasets/JuDDGES/pl-court-raw-enriched.Supreme-Court-Cases-1830-2019
US Supreme Court Legal Corpus (1830–2019)
Overview
A comprehensive, production-ready AI training dataset containing 456,589 documents from 122,930 US Supreme Court cases spanning 190 years (1830–2019).
This corpus captures the full adversarial record — petitions for certiorari, respondent briefs, reply briefs, amicus curiae filings, appendices, oral argument transcripts, and opinions. It is one of the most complete collections of Supreme Court procedural and… See the full description on the dataset page: https://huggingface.co/datasets/OwnedByDanes/Supreme-Court-Cases-1830-2019.courtlistenercases4court_opinions_filtered_under_25ksf_criminal_court
San Francisco Criminal Court Data
Linked criminal-court records for San Francisco County, combining scraped Superior Court docket data, District Attorney open-data feeds, and a charge-disposition spreadsheet that the Court produced only after sustained pressure under California Rules of Court rule 10.500.
Tables
File
Rows
Description
cases.parquet
77,406
Case-level records: case number, case ID, defendant name, filing date
register_of_actions.parquet
776,728… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/sf_criminal_court.court_opinions_filtered_full_sizeafrica-synth-governance-court-case-backlogs-all
African Court Case Backlogs | Africa (World Bank)
Size category: 10K<n<100K - Formats: csv - Sector: governance_security - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help analysts inspect structured… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-governance-court-case-backlogs-all.sf_criminal_court
San Francisco Criminal Court Data
Linked criminal-court records for San Francisco County, combining scraped Superior Court docket data, District Attorney open-data feeds, and a charge-disposition spreadsheet that the Court produced only after sustained pressure under California Rules of Court rule 10.500.
Tables
File
Rows
Description
cases.parquet
77,406
Case-level records: case number, case ID, defendant name, filing date
register_of_actions.parquet… See the full description on the dataset page: https://huggingface.co/datasets/Mateuzaoooo/sf_criminal_court.lower_court_insertion_swiss_judgment_predictionThis dataset contains an implementation of lower court insertion for the SwissJudgmentPrediction task.supreme_court_summary_entailment
Dataset Card for "supreme_court_summary_entailment"
More Information needed
lamus-roberts-court-legal-arguments
LAMUS: Roberts Court Legal Arguments (2005-2025)
The Current Supreme Court Era - Chief Justice John Roberts
📋 Dataset Description
This dataset contains 362,891 sentences from U.S. Supreme Court opinions during the Roberts Court era (2005-2025), automatically labeled with legal argument categories. This represents the current Supreme Court under Chief Justice John G. Roberts Jr.
Why Roberts Court?
The Roberts Court is particularly significant for… See the full description on the dataset page: https://huggingface.co/datasets/LavanyaPobbathi/lamus-roberts-court-legal-arguments.lk-appeal-court-judgements-chunkscalifornia_tos_court_cases_32k_v1courtThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "lekiwi_client",
"total_episodes": 20,
"total_frames": 17981,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lanhe1124/court.indian_supreme_court_judgements_en_ta
Indian Supreme Court Judgements Dataset (Sentence-Level, Translated to Tamil)
Overview
This dataset contains Indian Supreme Court judgements that have been split into sentences and translated into Tamil. The original judgements were sourced from the Indian Kanoon website. The dataset is useful for legal text processing, multilingual NLP tasks, and cross-lingual legal studies.
Data Processing Pipeline
Sentence Splitting:
Used pySBD (Python Sentence Boundary… See the full description on the dataset page: https://huggingface.co/datasets/Narenameme/indian_supreme_court_judgements_en_ta.us-public-tennis-courts
US Public Tennis Courts — 4,298 Locations Across 17 Metros (2026)
Geocoded public tennis court locations for 17 major US metros (Austin, Dallas-Fort Worth, Denver, Fort Lauderdale, Houston, Los Angeles, Miami, Orange County, Phoenix, Portland, San Diego, San Francisco, San Jose, Seattle, Tampa, Washington DC, West Palm Beach): 4,298 locations with latitude/longitude, court counts (11,527 individual courts where counted), lighting, practice walls, and surface where known. Derived… See the full description on the dataset page: https://huggingface.co/datasets/rrhagentbiz/us-public-tennis-courts.tos_court_opinions_filtered
