datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-india-law
Open India Law
Open, structured Indian primary law - plus the scrapers that build it.
Every judgment of the Supreme Court of India and all 25 High Courts, the decisions of 15
tribunals and regulators, and Central, State and Union Territory legislation down to the
individual section. Normalized to one schema, exclusively from official government sources.
Volume
Period
Court judgments
12,848,644
1950 to 2025
Tribunal and regulator matters
813,168
1985 to 2026… See the full description on the dataset page: https://huggingface.co/datasets/vaquill/open-india-law.Indian-Supreme-Court-Judgements-Chunked
Indian Supreme Court Judgements Chunked
Executive Summary
The dataset aims to address the chronic backlog in the Indian judiciary system, particularly in the Supreme Court, by creating a dataset optimized for legal language models (LLMs). The dataset will consist of pre-processed, chunked, and embedded textual data derived from the Supreme Court's judgment PDFs.
Problem and Importance - Motivation
Indian courts are overwhelmed with pending cases, with the… See the full description on the dataset page: https://huggingface.co/datasets/vihaannnn/Indian-Supreme-Court-Judgements-Chunked.indian-exams-rawdata
Indian Competitive Exams Raw Dataset (GATE, JEE Main & JEE Advanced)
This dataset contains raw PDF question papers, official answer keys, and extracted/parsed question structured data for major Indian national-level competitive engineering examinations: GATE, JEE Advanced, and JEE Main, spanning multiple years (2007–2025).
Data Sources & Attribution
The data in this repository was scraped and compiled from official conducting authority portals and public… See the full description on the dataset page: https://huggingface.co/datasets/Abhay557/indian-exams-rawdata.indian-legal-sections-bns-bnss-bsa-2023
🏛️ Indian Legal Sections — BNS · BNSS · BSA 2023
The First Structured, Unified JSON Dataset of Modern Indian Criminal Law
📖 Dataset Summary
This dataset contains 1,059 fully structured and verified sections extracted, parsed, and unified from India's three landmark criminal justice reform acts passed in December 2023. These three acts together replaced the colonial-era Indian Penal Code (IPC, 1860), the Code of Criminal Procedure… See the full description on the dataset page: https://huggingface.co/datasets/GSMS-B/indian-legal-sections-bns-bnss-bsa-2023.indian-legal-documents
Indian Legal Documents
Open Indian statutory and legal-document data for AI, search, and legal research.
This dataset is part of the KanoonGPT Open Legal Data Initiative — an effort to make Indian legal data easier to access, structure, and build on for open-source research, legal tech, and production AI systems.
KanoonGPT is building structured Indian legal datasets and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.… See the full description on the dataset page: https://huggingface.co/datasets/KanoonGPT/indian-legal-documents.Chunked-Indian-Supreme-Court-Judgements
Indian Supreme Court Judgements Chunked
india-acts
India Acts — Central & State Statutes
A comprehensive corpus of Indian legislation in PDF form — covering both Central (Parliament) Acts and State / Union Territory Acts — in English and Hindi, scraped and consolidated from publicly available government sources (primarily the India Code portal and individual State legislature websites).
This dataset is intended as a research and AI-training resource for tasks such as legal document retrieval, statutory question-answering… See the full description on the dataset page: https://huggingface.co/datasets/judicialmind/india-acts.indian-government-schemes-2025
Indian Government Schemes Dataset 2026
Dataset Description
The most comprehensive structured dataset of Indian central and state government schemes — 4,693 schemes across all ministries and states, with machine-readable eligibility fields.
Maintained by SmartDuke Technologies · Coimbatore, Tamil Nadu, India
This dataset powers SchemeFit — India's government scheme finder for citizens and businesses.
What Makes This Different
Most existing Indian… See the full description on the dataset page: https://huggingface.co/datasets/smartduketech/indian-government-schemes-2025.Indian_Supreme_Court_Judgments
Indian Supreme Court Judgments (Structured JSON)
About Dataset
This dataset contains fully extracted, structured JSON data for Indian Supreme Court judgments. This is a parsed, machine-readable version of the raw PDF repository, designed specifically for Natural Language Processing (NLP), RAG (Retrieval-Augmented Generation), and legal tech machine learning applications.
The dataset is provided in .jsonl (JSON Lines) format. Each row represents a single case and contains… See the full description on the dataset page: https://huggingface.co/datasets/sinhal/Indian_Supreme_Court_Judgments.indian-history-hindi-QA-3.4k
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
This dataset contains 3.47k top-notch question-answer pairs about Indian History in Hindi.
Curated by: Mohd Kaif
Language(s) (NLP): Hindi
License: apache-2.0
Indian-Legal-QA-BNS-BNSS-BSA
Indian Legal QA — BNS + BNSS + BSA 2023
6,354 structured question-answer pairs covering all 1,059 sections across India's three criminal justice acts of 2023
Overview
This dataset contains 6,354 instruction-format question-answer pairs in JSONL format, covering every section of India's three criminal justice reform acts enacted in 2023. Each section has exactly 6 questions approaching the same legal provision from different angles… See the full description on the dataset page: https://huggingface.co/datasets/GSMS-B/Indian-Legal-QA-BNS-BNSS-BSA.indian_law
Indian Law Dataset
The Indian Law Dataset is a high-quality, open-source dataset (~50M tokens) focused on Indian jurisprudence. It provides structured chain-of-thought reasoning traces across 10+ branches of law, enabling the training and evaluation of advanced reasoning-capable language models.
Summary
• Domain: Law / Indian Jurisprudence / Legal Reasoning
• Scale: ~50M tokens, 47,789 rows
• Source: Generated with advanced distillation techniques using structured… See the full description on the dataset page: https://huggingface.co/datasets/169Pi/indian_law.indian-legal-exam-benchmark
Indian Law Entrance Exam Papers Dataset
A comprehensive, research-ready collection of multiple-choice questions from key Indian law entrance exams, including CLAT (UG & PG) and Delhi Judicial Service papers. This dataset is ideal for legal education, NLP benchmarking, and educational technology.
Paper: Are LLMs Court-Ready? Evaluating Frontier Models on Indian Legal Reasoning
Overview
This dataset provides 6,218 exam questions from 38 law entrance papers conducted… See the full description on the dataset page: https://huggingface.co/datasets/adalat-ai/indian-legal-exam-benchmark.Dataset-For-Indian-legal-knowledge-base About This Dataset
This dataset is the knowledge backbone of LegalEagle — an AI-powered contract review platform for Indian startups and freelancers. It contains Indian statutes, contract templates, landmark case references, and clause examples, curated specifically for retrieval-augmented generation (RAG) in the Indian legal domain.
All government statutes included are in the public domain (Government of India publications).
Dataset Structure
dataset/
├── acts/… See the full description on the dataset page: https://huggingface.co/datasets/d-riti/Dataset-For-Indian-legal-knowledge-base.k12-indian-curriculum-4.9m
BharatLLM K-12 Indian Curriculum Dataset (4.9M)
4,904,936 question-answer pairs covering CBSE/NCERT K-12 curriculum across 12 Indian languages.
Language
Script
Entries
English
Latin
~594K
Hindi
Devanagari
~449K
Bengali
Bengali
~408K
Telugu
Telugu
~408K
Tamil
Tamil
~408K
Kannada
Kannada
~408K
Malayalam
Malayalam
~408K
Marathi
Devanagari
~408K
Gujarati
Gujarati
~408K
Odia
Odia
~408K
PunjabiGurmukhi
~408K
Urdu
Nastaliq
~374K
Format
{… See the full description on the dataset page: https://huggingface.co/datasets/FoundryAILabs/k12-indian-curriculum-4.9m.IndiaFinBench
IndiaFinBench
An Evaluation Benchmark for Large Language Model Performance on Indian Financial Regulatory Text
Rajveer Singh Pall · Gyan Ganga Institute of Technology and Sciences, Jabalpur, India
406 QA Pairs
192 Source Documents
4 Task Types
12 Models Evaluated
Expert-annotated
SEBI · RBI · 1992–2026
REG · NUM · CON · TMP
Zero-shot, full benchmark
Overview
IndiaFinBench is, to our knowledge, the first publicly available evaluation… See the full description on the dataset page: https://huggingface.co/datasets/Rajveer-code/IndiaFinBench.k12-indian-curriculum-4.9m
BharatLLM K-12 Indian Curriculum Dataset (4.9M)
4,904,936 question-answer pairs covering CBSE/NCERT K-12 curriculum across 12 Indian languages.
Language
Script
Entries
English
Latin
~594K
Hindi
Devanagari
~449K
Bengali
Bengali
~408K
Telugu
Telugu
~408K
Tamil
Tamil
~408K
Kannada
Kannada
~408K
Malayalam
Malayalam
~408K
Marathi
Devanagari
~408K
Gujarati
Gujarati
~408K
Odia
Odia
~408K
Punjabi
Gurmukhi
~408K
Urdu
Nastaliq
~374K
Format… See the full description on the dataset page: https://huggingface.co/datasets/krittus/k12-indian-curriculum-4.9m.Indian_Penal_Code
Indian Penal Code Dataset
Dataset Description:
The Indian Penal Code (IPC) Book PDF presents a rich and comprehensive dataset that holds immense potential for advancing Natural Language Processing (NLP) tasks and Language Model applications. This dataset encapsulates the entire spectrum of India's criminal law, offering a diverse range of legal principles, provisions, and case laws. With its intricate language and multifaceted legal content, the IPC dataset provides a… See the full description on the dataset page: https://huggingface.co/datasets/harshitv804/Indian_Penal_Code.Indian-legal-data-v3
Indian Legal Dataset V3
Overview
Indian Legal Dataset V3 is a large-scale instruction-tuning dataset focused on Indian law, constitutional law, criminal law, legal reasoning, legal drafting, and real-world legal assistance.
Compared to V2, this version expands the dataset with:
legal drafting instruction pairs,
hypothetical legal scenarios,
detailed IPC-focused data,
practical real-world legal instructions,
concise legal QA pairs.
After integrating the new data sources… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/Indian-legal-data-v3.indian-pharma-dataset-2026-augast
PharmaLens: 200k Medicine Catalog (Salts, Prices, Interactions and Reviews)
I spent weeks compiling and cleaning this retrieval database for a project. Instead of letting 300MB+ of structured pharmaceutical data sit idle on my hard drive, I am open-sourcing it. Use it for your RAG pipelines, chatbots, pricing tools, or whatever else you are building.
Overview
Finding clean, structured pharmaceutical datasets with commercial brand names, active salt compositions… See the full description on the dataset page: https://huggingface.co/datasets/sinhal/indian-pharma-dataset-2026-augast.Indian-Income-Tax-Returns
Indian Income Tax Return Synthetic Dataset
A fully synthetic, high-fidelity dataset of Indian Income Tax Return forms (ITR-4, ITR-5, and ITR-6). Designed to support OCR, text extraction, table parsing, tax attribute detection, document intelligence, and LLM fine-tuning for structured data extraction. Each record includes a PDF tax return and a matching structured JSON file containing parsed fields.
This dataset simulates realistic taxpayer filings across:
Individuals (with Aadhaar… See the full description on the dataset page: https://huggingface.co/datasets/AgamiAI/Indian-Income-Tax-Returns.indian-legal-opposing-counsel-dataset
⚖️ Indian Legal Opposing Counsel Dataset
A combined, preprocessed dataset of 26,326 examples for training an Indian legal opposing counsel AI model. Ready-to-use in ChatML format for SFT training.
📊 Dataset Stats
Split
Rows
Size
Train
25,009
65 MB
Test
1,317
3.5 MB
Total
26,326
69 MB
📦 Sources
Source Dataset
Rows
Content
viber1/indian-law-dataset
24,607
Writs, PIL, civil procedure, constitutional law, IPC… See the full description on the dataset page: https://huggingface.co/datasets/pkheria7/indian-legal-opposing-counsel-dataset.Indian-Legal-SFT-Dataset
Vidhaan: High-Density Indian Legal Instruction Dataset
Vidhaan is a comprehensive, high-precision instruction-tuning dataset containing 20,690 QA pairs derived from 113 Central Acts of India. It was built specifically to solve the "context-splitting" problem found in standard legal RAG datasets.
🛠 Dataset Structure & Format
Primary File: vidhaan_training_v1.jsonl
Format: JSON Lines (JSONL)
Schema: - instruction: (String) A precise legal query.
context: (String) The… See the full description on the dataset page: https://huggingface.co/datasets/SharathReddy/Indian-Legal-SFT-Dataset.Indian-legal-data-v2
Legal Instruction Dataset (v2)
📌 Overview
This dataset contains high-quality instruction–response pairs derived from Indian legal texts, primarily focusing on statutory interpretation and structured legal explanations.
Version 2 represents a significant scale and quality upgrade over v1:
v1: 33,077 samples
v2: 171,640 samples
The dataset is designed specifically for instruction tuning of language models, emphasizing clarity, structure, and legal reasoning patterns.… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/Indian-legal-data-v2.indian-legal-records
LH2 Data — Indian Legal Records & Judgments Corpus
The most comprehensive structured Indian legal records corpus available for AI training — 267M+ case records spanning the full judicial hierarchy, paired with a pre-computed AI enrichment layer across 21M+ court orders.
Dataset Summary
This corpus provides structured, indexed, and partially labelled legal records from the Indian judicial system at a scale that has no public equivalent. It covers the Supreme Court of… See the full description on the dataset page: https://huggingface.co/datasets/LH2-data-labs/indian-legal-records.Indian-legal-data-v1
Legal Instruction Dataset
📌 Overview
This dataset contains instruction–response pairs derived from sections of the Indian Acts.
The dataset is designed for instruction tuning of language models, with a focus on:
structured legal explanations
bullet-point formatting
long-form responses
🧠 Dataset Description
Task Type: Instruction Tuning / Legal QA
Domain: Indian Law
Language: English
Format: JSONL
🔥 What makes this good (not generic fluff)… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/Indian-legal-data-v1.indian_protocols_based_clinical_QnA
Indian Protocols-Based Clinical Q&A
A rubric-graded evaluation dataset built from clinical guideline documents (Indian and international). Each sample is a realistic doctor-side query against a known protocol, paired with rubrics that grade (a) whether the system retrieved/identified the correct guideline content and (b) whether the final answer is clinically complete and safe.
What this evaluates
This dataset is built to stress-test clinical assistants on… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/indian_protocols_based_clinical_QnA.indian-agri-advice-multilingual
This dataset is a quality-filtered golden subset prepared using a layered regex + local LLM-as-judge pipeline, then re-adapted using Adaption's Adaptive Data platform.
Indian Agricultural Advisory Dataset — Multilingual (Golden v5)
718 rows | 11 languages | 14 agro-climatic zones | 12 categories | 100% metadata fill rate
A multilingual agricultural advisory dataset covering 14 of India's 15 Planning Commission agro-climatic zones, localized to 11 Indian languages. This is a… See the full description on the dataset page: https://huggingface.co/datasets/vinod-anbalagan/indian-agri-advice-multilingual.india-state-debt-instruction-dataset
🇮🇳 India State Debt Chain-of-Thought (CoT) Instruction Dataset (2020–2026)
This dataset contains 761 high-quality instruction-following prompt-response pairs focusing on public debt liabilities, Debt-to-GSDP percentages, per-capita debt, fiscal vulnerability classifications, and multi-year growth trajectories across 31 Indian States and Union Territories (2020–2026).
Data Schema & Format
Every record includes structured Chain-of-Thought (CoT) reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Sushilmishrawork/india-state-debt-instruction-dataset.study-in-india-faq
Study-in-India FAQ Dataset
This dataset, study-in-india-faq, is designed for fine-tuning language models to answer frequently asked questions about studying in India. It includes 200,000 pairs of questions and answers covering topics such as admissions, scholarships, accommodation, cultural adjustments, and visa requirements.
Dataset Summary
The study-in-india-faq dataset is a comprehensive resource for students seeking information about studying in India, whether they… See the full description on the dataset page: https://huggingface.co/datasets/millat/study-in-india-faq.
