datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Dataset-For-Indian-legal-knowledge-base About This Dataset
This dataset is the knowledge backbone of LegalEagle — an AI-powered contract review platform for Indian startups and freelancers. It contains Indian statutes, contract templates, landmark case references, and clause examples, curated specifically for retrieval-augmented generation (RAG) in the Indian legal domain.
All government statutes included are in the public domain (Government of India publications).
Dataset Structure
dataset/
├── acts/… See the full description on the dataset page: https://huggingface.co/datasets/d-riti/Dataset-For-Indian-legal-knowledge-base.Indian-legal-data-v3
Indian Legal Dataset V3
Overview
Indian Legal Dataset V3 is a large-scale instruction-tuning dataset focused on Indian law, constitutional law, criminal law, legal reasoning, legal drafting, and real-world legal assistance.
Compared to V2, this version expands the dataset with:
legal drafting instruction pairs,
hypothetical legal scenarios,
detailed IPC-focused data,
practical real-world legal instructions,
concise legal QA pairs.
After integrating the new data sources… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/Indian-legal-data-v3.indian-pharma-dataset-2026-augast
PharmaLens: 200k Medicine Catalog (Salts, Prices, Interactions and Reviews)
I spent weeks compiling and cleaning this retrieval database for a project. Instead of letting 300MB+ of structured pharmaceutical data sit idle on my hard drive, I am open-sourcing it. Use it for your RAG pipelines, chatbots, pricing tools, or whatever else you are building.
Overview
Finding clean, structured pharmaceutical datasets with commercial brand names, active salt compositions… See the full description on the dataset page: https://huggingface.co/datasets/sinhal/indian-pharma-dataset-2026-augast.indian-legal-opposing-counsel-dataset
⚖️ Indian Legal Opposing Counsel Dataset
A combined, preprocessed dataset of 26,326 examples for training an Indian legal opposing counsel AI model. Ready-to-use in ChatML format for SFT training.
📊 Dataset Stats
Split
Rows
Size
Train
25,009
65 MB
Test
1,317
3.5 MB
Total
26,326
69 MB
📦 Sources
Source Dataset
Rows
Content
viber1/indian-law-dataset
24,607
Writs, PIL, civil procedure, constitutional law, IPC… See the full description on the dataset page: https://huggingface.co/datasets/pkheria7/indian-legal-opposing-counsel-dataset.Indian-Legal-SFT-Dataset
Vidhaan: High-Density Indian Legal Instruction Dataset
Vidhaan is a comprehensive, high-precision instruction-tuning dataset containing 20,690 QA pairs derived from 113 Central Acts of India. It was built specifically to solve the "context-splitting" problem found in standard legal RAG datasets.
🛠 Dataset Structure & Format
Primary File: vidhaan_training_v1.jsonl
Format: JSON Lines (JSONL)
Schema: - instruction: (String) A precise legal query.
context: (String) The… See the full description on the dataset page: https://huggingface.co/datasets/SharathReddy/Indian-Legal-SFT-Dataset.indian-legal-records
LH2 Data — Indian Legal Records & Judgments Corpus
The most comprehensive structured Indian legal records corpus available for AI training — 267M+ case records spanning the full judicial hierarchy, paired with a pre-computed AI enrichment layer across 21M+ court orders.
Dataset Summary
This corpus provides structured, indexed, and partially labelled legal records from the Indian judicial system at a scale that has no public equivalent. It covers the Supreme Court of… See the full description on the dataset page: https://huggingface.co/datasets/LH2-data-labs/indian-legal-records.Indian-legal-data-v2
Legal Instruction Dataset (v2)
📌 Overview
This dataset contains high-quality instruction–response pairs derived from Indian legal texts, primarily focusing on statutory interpretation and structured legal explanations.
Version 2 represents a significant scale and quality upgrade over v1:
v1: 33,077 samples
v2: 171,640 samples
The dataset is designed specifically for instruction tuning of language models, emphasizing clarity, structure, and legal reasoning patterns.… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/Indian-legal-data-v2.Indian-legal-data-v1
Legal Instruction Dataset
📌 Overview
This dataset contains instruction–response pairs derived from sections of the Indian Acts.
The dataset is designed for instruction tuning of language models, with a focus on:
structured legal explanations
bullet-point formatting
long-form responses
🧠 Dataset Description
Task Type: Instruction Tuning / Legal QA
Domain: Indian Law
Language: English
Format: JSONL
🔥 What makes this good (not generic fluff)… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/Indian-legal-data-v1.indian-legal-dataset-indian-law
Indian Legal Q&A Dataset (Merged)
This dataset contains approximately 1,500 fine-tuning pairs for the Indian legal domain. It consists of instruction, input (context), and output sequences, covering various facets of Indian legislation.
Content Summary
The dataset is a consolidated collection of Q&A pairs covering:
The Constitution of India
Indian Penal Code (IPC)
Indian Contract Act, 1872
Right to Information (RTI) Act, 2005
Criminal Procedure Code (CrPC)
Civil… See the full description on the dataset page: https://huggingface.co/datasets/RMani1/indian-legal-dataset-indian-law.indian-family-law-data
Indian Family Law QA Dataset
This dataset contains 1,401 high-quality question-answer pairs focused on Family Law II, specifically tailored for legal studies in the Indian context. The dataset covers critical areas such as Hindu Succession, Coparcenary rights, and principles of Muslim Law.
Dataset Details
Total Rows: 1,401
Language: English
Task: Question Answering / Legal Knowledge Retrieval
Format: CSV (two-column format)
Dataset Structure
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/steamed-potatop/indian-family-law-data.indian-legal-sft-data-part-1
Dataset Card for indian-legal-sft-data-part-1
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/Prarabdha/indian-legal-sft-data-part-1/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/Prarabdha/indian-legal-sft-data-part-1.indian-legal-data-sft1.1
Dataset Card for indian-legal-data-sft1.1
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/Prarabdha/indian-legal-data-sft1.1/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/Prarabdha/indian-legal-data-sft1.1.indian_constitution_dataindian-legal-data-sft-part-2
Dataset Card for indian-legal-data-sft-part-2
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/Prarabdha/indian-legal-data-sft-part-2/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/Prarabdha/indian-legal-data-sft-part-2.indian-legal-data-sft1.2
Dataset Card for indian-legal-data-sft1.2
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/Prarabdha/indian-legal-data-sft1.2/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/Prarabdha/indian-legal-data-sft1.2.indian-legal-data-sft1.3
Dataset Card for indian-legal-data-sft1.3
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/Prarabdha/indian-legal-data-sft1.3/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/Prarabdha/indian-legal-data-sft1.3.
