CoolFace
Datasetpublic

Shreyasrao/Indian-law-supreme-court-judgements-2016

Dataset Card: Indian Supreme Court Judgments 2016 Dataset Description A comprehensive, structured dataset of 589 judgments delivered by the Supreme Court of India during the calendar year 2016, plus a small number of spill-over judgments from 2017 bundled in the original source PDFs. Each judgment is available in three forms: the original scanned PDF as downloaded from the official eCourts portal, a full text extracted via OCR as Markdown, and a fully structured… See the full description on the dataset page: https://huggingface.co/datasets/Shreyasrao/Indian-law-supreme-court-judgements-2016.

sourceHugging Facecc-by-4.0updated 3mo agoView on Hugging Face
2likes144downloads
Dataset Card

Dataset Card: Indian Supreme Court Judgments 2016

Dataset Description

A comprehensive, structured dataset of 589 judgments delivered by the Supreme Court of India during the calendar year 2016, plus a small number of spill-over judgments from 2017 bundled in the original source PDFs. Each judgment is available in three forms: the original scanned PDF as downloaded from the official eCourts portal, a full text extracted via OCR as Markdown, and a fully structured JSON with extracted entities (case title, summary, judges, parties, legal sections, and topic labels across 11 categories).

The dataset was curated by downloading the 2016 subset from the Indian Supreme Court Judgments registry on AWS Open Data (managed by Dattam Labs, CC-BY-4.0), then OCR-processing every PDF through PaddleOCR-VL-1.6 to produce the Markdown and structured JSON outputs.

Dataset Structure

indian-supreme-court-judgements-2016/
  raw_data/          # 589 PDFs (282 MB compressed, 368 MB uncompressed)
  extracted_mds/     # 589 Markdown files (28 MB) — full OCR text
  extracted_jsons/   # 589 JSON files (4.1 MB) — structured entities
  DATASET_CARD.md    # this file

Naming Convention

Files follow the pattern {year}_{month}_{page_start}_{page_end}_{language}.{ext}, where:

  • —year: 2016 (one 2017 spill-over bundle from same source)
  • —month: numeric month of delivery (1=January through 12=December)
  • —page_start / page_end: page range within the original eCourts pagination scheme
  • —language: EN (English only in this dataset)
  • —ext: pdf / md / json

Example: 2016-1-179-186-en.json → judgment delivered in January 2016 (month 1), spanning pages 179-186 of the eCourts index, in English.

Data Formats

1. Raw PDFs (raw_data/)

Original scanned PDFs from the eCourts portal. 589 files, 368 MB total. Contains the full scanned judgment text including OCR layer where available.

2. Extracted Markdown (extracted_mds/)

Full text extracted via PaddleOCR-VL-1.6, formatted as Markdown with:

  • —A YAML-style header with metadata (page count, processing time)
  • —Full OCR output per page with ## Page N separators
  • —Case law references, citations, and the complete judgment text
  • —Error annotations where OCR processing of a specific page failed (shown inline)

Example header:

markdown
# 2016_1_179_186_EN
*Converted via PaddleOCR-VL-1.6 | 8 pages | 37s*
3. Structured JSON (extracted_jsons/)

Machine-readable JSON with the following schema:

json
{
  "filename": "2016-1-179-186-en.md",
  "metadata": {
    "year": 2016,
    "month": 1,
    "page_start": 179,
    "page_end": 186,
    "language": "en"
  },
  "entities": {
    "case_title": { "title": "State of Assam v. Ramen Dowarah" },
    "summary": { "summary": "The Supreme Court restored..." },
    "judges": [{ "name": "A. N. Joseph", "role": "bench_member" },
               { "name": "Arun Mishra",   "role": "author" }],
    "parties": [{ "name": "State of Assam", "role": "appellant" },
                { "name": "Ramen Dowarah",  "role": "respondent" }],
    "sections": [{ "section": "302",              "act": "Indian Penal Code" },
                 { "section": "376",              "act": "Indian Penal Code" },
                 { "section": "164",              "act": "Code of Criminal Procedure" }],
    "topics": [{ "text": "Rape",                  "category": "criminal" },
               { "text": "Murder",                "category": "criminal" },
               { "text": "Dying Declaration",     "category": "procedural" }]
  },
  "raw_text_preview": "### STATE OF ASSAM RAMEN DOWARAH..."
}

Dataset Statistics

MetricValue
Total judgments589
Total pages11,667
Raw PDF size367.7 MB
Extracted Markdown size26.2 MB
Structured JSON size4.1 MB
JSON size (per file)2.2–8.9 KB (avg 7.3 KB)
Cases with summary586 / 589
Cases with judges extracted582 / 589
Cases with parties extracted585 / 589
Cases with title584 / 589
Total section references2,765
Total topic labels3,041 (2,291 unique)
Unique acts referenced447
Unique judges34

Monthly Distribution

MonthCasesMonthCases
January80July35
February108August51
March68September50
April66October1
May58November36
June24December12

The spike in February (108 cases) and concentration in the first half of the year reflects the Court's calendar — the Supreme Court typically hears its heaviest docket during the pre-summer term.

Topic Category Distribution

CategoryLabelsDescription
Procedural840Criminal appeals, writs, evidentiary issues, jurisdiction
Administrative511Service law, disciplinary proceedings, administrative orders
Criminal455Murder, rape, NDPS, corruption, criminal appeals
Constitutional346Fundamental rights, constitutional interpretation, ordinances
Civil294Contract, tort, tenancy, civil procedure
Property194Land acquisition, lease, eviction, title disputes
Tax159Income tax, customs, excise, GST
Corporate106Company law, winding up, SARFAESI, insolvency
Other70Miscellaneous
Family46Divorce, maintenance, succession, custody
Environmental20Forest conservation, pollution, mining

Top Legal Topics

The most frequently occurring topics include Statutory Interpretation (42 cases), Murder (38), Service Law (34), Land Acquisition (24), Criminal Appeal (23), Public Interest Litigation (21), Compensation (14), Judicial Review (14), Arbitration (14), and Writ Jurisdiction (14).

Most Frequently Cited Acts

ActReferences
Indian Penal Code437
Constitution of India386
Land Acquisition Act, 189499
Code of Criminal Procedure, 197397 (combined with "Code of Criminal Procedure": 186)
Income Tax Act, 196176
Code of Civil Procedure, 190871
Arbitration and Conciliation Act, 199658
Indian Evidence Act38
Electricity Act, 200332
Prevention of Corruption Act, 198828

Most Frequent Judges

JudgeCases
Dipak Misra104
T. S. Thakur81
A. K. Sikri74
Rohinton F. Nariman74
Kurian Joseph68
Shiva Kirti Singh66
Abhay Manohar Sapre63
R. Banumathi57
Prafulla C. Pant53
V. Gopala Gowda42

Data Provenance

  1. 1.Source: Indian Supreme Court Judgments on AWS Open Data Registry, managed by Dattam Labs. The complete dataset covers judgments from 1950 to 2025 in both English and regional Indian languages. Only the 2016 English subset was used here.
  1. 1.Download date: 27 May 2026, via aws s3 cp --no-sign-request.
  1. 1.OCR processing: Every PDF was processed through PaddleOCR-VL-1.6 (a vision-language OCR model by PaddlePaddle), running via llama.cpp server. Each judgment took 10–60 seconds depending on page count. The OCR output was captured as Markdown with per-page pagination.
  1. 1.Entity extraction: A structured JSON was generated from each Markdown file by LLM-based information extraction, identifying: case title, summary, judges on the bench (with author designation), parties with roles, legal sections cited (with corresponding act names), and topic labels (text + category).
  1. 1.Topic classification: Topics are assigned into 11 broad categories: procedural, administrative, criminal, constitutional, civil, property, tax, corporate, other, family, and environmental.

Data Quality Notes

  • —OCR fidelity: PaddleOCR-VL-1.6 is a vision-language model, not a traditional OCR engine. Text extraction quality is high for clean scanned pages but may exhibit errors on:
  • —Faint or degraded scans (common in older judgments)
  • —Complex layouts (multi-column text, footnotes)
  • —Tables and non-standard formatting
  • —Handwritten annotations

Some pages failed OCR entirely and are marked with:

  [Error processing page N: Exception from the 'vlm' worker: ...]
  • —Summary generation: 3 cases lack AI-generated summaries (likely very short or header-only PDFs). 5 case titles could not be extracted. 7 judgments have no judge information and 4 have no party information — these are very short procedural orders where entity extraction was unreliable.
  • —Section deduplication: The same act may appear under slightly different names (e.g., "Code of Criminal Procedure, 1973" vs "Code of Criminal Procedure" vs "Cr.P.C."). No normalization has been applied; all forms appear as extracted.
  • —Topic sparsity: The 3,041 topic labels span 2,291 unique topics, indicating high diversity. Only Statutory Interpretation (42) and Murder (38) appear with significant frequency; most topics are unique to a single case. This reflects the wide-ranging docket of the Supreme Court.
  • —The October anomaly: Only 1 judgment is recorded for month 10 (October). This is not a gap — the Supreme Court of India typically observes a long autumn/Diwali vacation in October, during which very few bench days occur.
  • —Spill-over 2017 file: One PDF (2017-5-160-276) is included in the source data despite being filed under the 2016 directory. It has been retained as-is. It is a 7-judge constitution bench judgment on the validity of Bihar's ordinance re-promulgation practice.

Intended Uses

This dataset is suitable for:

  • —Legal NLP research: Judgment summarization, legal text classification, topic modelling, and named entity recognition on Indian case law
  • —Legal information retrieval: Training query-based retrieval systems for Indian Supreme Court jurisprudence
  • —Legal analytics: Studying court trends, judicial behaviour, citation networks, and topic evolution
  • —Legal education: Providing structured, labelled access to landmark and routine Supreme Court judgments for curriculum development
  • —Benchmarking OCR systems: The raw PDFs paired with cleaned structured JSON serve as an evaluation set for legal domain OCR

Limitations and Biases

  • —Single year: This dataset covers only 2016 (plus one 2017 bundle). It does not capture temporal trends or the evolution of legal doctrine across decades.
  • —English only: Only English-language judgments are included. The source dataset includes judgments in Hindi and other regional languages which are excluded here.
  • —Topic taxonomy: The 11-category topic model was chosen for broad coverage, but it inevitably loses nuance at the boundaries (e.g., certain administrative matters touch on constitutional questions).
  • —OCR quality variance: Older, poorly scanned PDFs produce lower quality OCR output, which in turn affects entity extraction reliability.
  • —Judge role assignment: The "author" vs "bench_member" distinction in the JSON is an LLM inference, not official court metadata. For multi-judge opinions the designated author may not always be correctly identified.

License

The underlying source data is licensed under CC-BY-4.0 as per the AWS Open Data Registry entry. The derived Markdown and JSON outputs in this dataset are made available under the same CC-BY-4.0 license.

Citation

If you use this dataset, please cite:

bibtex
@misc{indian_supreme_court_judgments_2016,
  title = {Indian Supreme Court Judgments 2016: A Structured Dataset},
  author = {Shreyas Rao},
  year = {2026},
  month = {May},
  howpublished = {\url{https://registry.opendata.aws/indian-supreme-court-judgments/}},
  note = {Accessed 27 May 2026. OCR-processed with PaddleOCR-VL-1.6.}
}