datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Indian-legal-data-v3
Indian Legal Dataset V3
Overview
Indian Legal Dataset V3 is a large-scale instruction-tuning dataset focused on Indian law, constitutional law, criminal law, legal reasoning, legal drafting, and real-world legal assistance.
Compared to V2, this version expands the dataset with:
legal drafting instruction pairs,
hypothetical legal scenarios,
detailed IPC-focused data,
practical real-world legal instructions,
concise legal QA pairs.
After integrating the new data sources… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/Indian-legal-data-v3.Indian-Legal-SFT-Dataset
Vidhaan: High-Density Indian Legal Instruction Dataset
Vidhaan is a comprehensive, high-precision instruction-tuning dataset containing 20,690 QA pairs derived from 113 Central Acts of India. It was built specifically to solve the "context-splitting" problem found in standard legal RAG datasets.
🛠 Dataset Structure & Format
Primary File: vidhaan_training_v1.jsonl
Format: JSON Lines (JSONL)
Schema: - instruction: (String) A precise legal query.
context: (String) The… See the full description on the dataset page: https://huggingface.co/datasets/SharathReddy/Indian-Legal-SFT-Dataset.Indian-legal-data-v2
Legal Instruction Dataset (v2)
📌 Overview
This dataset contains high-quality instruction–response pairs derived from Indian legal texts, primarily focusing on statutory interpretation and structured legal explanations.
Version 2 represents a significant scale and quality upgrade over v1:
v1: 33,077 samples
v2: 171,640 samples
The dataset is designed specifically for instruction tuning of language models, emphasizing clarity, structure, and legal reasoning patterns.… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/Indian-legal-data-v2.Indian_Climate_Disaster_data
BharatCRIC: Indian Climate Disaster Data (Heatwave Advisories + Scam Pairs)
Generated by scripts/grade_a_rebuild.py with blueprint grade_a_2026_04_30.
File
Use
instruction_dataset_main.jsonl
Full structured instruction upload (1680 rows)
instruction_dataset_smoke.jsonl
50-row smoke test, 5 languages x 5 formats x 2 rows
instruction_dataset_smoke_v2.jsonl
Same smoke set for the previous filename expected by notes
preference_pairs_scams.jsonl
Mirrored genuine-vs-scam… See the full description on the dataset page: https://huggingface.co/datasets/sahilmaniyar888/Indian_Climate_Disaster_data.indian-farmer-negotiation-data
🌾 Indian Farmer Mandi Negotiation Dataset
A high-quality, realistic training dataset for building AI systems that help Indian farmers negotiate better prices with traders at mandis (agricultural markets).
Dataset Details
Size: 5,000 examples
Language: Hindi / Hinglish (natural spoken style)
Coverage: 30 crops × 18 Indian states
Format: Input–Output pairs for supervised fine-tuning
Input Fields
Each example's input contains:
Field
Description
Example… See the full description on the dataset page: https://huggingface.co/datasets/StackOverflowed512/indian-farmer-negotiation-data.
