datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
minuszero-indian-autonomous-driving-dataset
Minus Zero Indian Urban Autonomous Driving Dataset
Overview
This dataset provides original multicamera autonomous-driving recordings in MCAP format. It is designed for non-commercial research on surround-view perception, temporal and cross-camera synchronization, H.265 video pipelines, localization, GNSS/pose integration, and robotics data tooling.
Recordings include camera and GNSS/pose streams, with machine-state telemetry present in a small subset. Camera… See the full description on the dataset page: https://huggingface.co/datasets/gagandeepreehal/minuszero-indian-autonomous-driving-dataset.indian-law-datasetindian-pharma-dataset-2026-augast
PharmaLens: 200k Medicine Catalog (Salts, Prices, Interactions and Reviews)
I spent weeks compiling and cleaning this retrieval database for a project. Instead of letting 300MB+ of structured pharmaceutical data sit idle on my hard drive, I am open-sourcing it. Use it for your RAG pipelines, chatbots, pricing tools, or whatever else you are building.
Overview
Finding clean, structured pharmaceutical datasets with commercial brand names, active salt compositions… See the full description on the dataset page: https://huggingface.co/datasets/sinhal/indian-pharma-dataset-2026-augast.Indian-legal-data-v3
Indian Legal Dataset V3
Overview
Indian Legal Dataset V3 is a large-scale instruction-tuning dataset focused on Indian law, constitutional law, criminal law, legal reasoning, legal drafting, and real-world legal assistance.
Compared to V2, this version expands the dataset with:
legal drafting instruction pairs,
hypothetical legal scenarios,
detailed IPC-focused data,
practical real-world legal instructions,
concise legal QA pairs.
After integrating the new data sources… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/Indian-legal-data-v3.Indian-Legal-SFT-Dataset
Vidhaan: High-Density Indian Legal Instruction Dataset
Vidhaan is a comprehensive, high-precision instruction-tuning dataset containing 20,690 QA pairs derived from 113 Central Acts of India. It was built specifically to solve the "context-splitting" problem found in standard legal RAG datasets.
🛠 Dataset Structure & Format
Primary File: vidhaan_training_v1.jsonl
Format: JSON Lines (JSONL)
Schema: - instruction: (String) A precise legal query.
context: (String) The… See the full description on the dataset page: https://huggingface.co/datasets/SharathReddy/Indian-Legal-SFT-Dataset.Indian-legal-data-v2
Legal Instruction Dataset (v2)
📌 Overview
This dataset contains high-quality instruction–response pairs derived from Indian legal texts, primarily focusing on statutory interpretation and structured legal explanations.
Version 2 represents a significant scale and quality upgrade over v1:
v1: 33,077 samples
v2: 171,640 samples
The dataset is designed specifically for instruction tuning of language models, emphasizing clarity, structure, and legal reasoning patterns.… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/Indian-legal-data-v2.Indian-legal-data-v1
Legal Instruction Dataset
📌 Overview
This dataset contains instruction–response pairs derived from sections of the Indian Acts.
The dataset is designed for instruction tuning of language models, with a focus on:
structured legal explanations
bullet-point formatting
long-form responses
🧠 Dataset Description
Task Type: Instruction Tuning / Legal QA
Domain: Indian Law
Language: English
Format: JSONL
🔥 What makes this good (not generic fluff)… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/Indian-legal-data-v1.indian-legal-dataset-indian-law
Indian Legal Q&A Dataset (Merged)
This dataset contains approximately 1,500 fine-tuning pairs for the Indian legal domain. It consists of instruction, input (context), and output sequences, covering various facets of Indian legislation.
Content Summary
The dataset is a consolidated collection of Q&A pairs covering:
The Constitution of India
Indian Penal Code (IPC)
Indian Contract Act, 1872
Right to Information (RTI) Act, 2005
Criminal Procedure Code (CrPC)
Civil… See the full description on the dataset page: https://huggingface.co/datasets/RMani1/indian-legal-dataset-indian-law.Indian_Climate_Disaster_data
BharatCRIC: Indian Climate Disaster Data (Heatwave Advisories + Scam Pairs)
Generated by scripts/grade_a_rebuild.py with blueprint grade_a_2026_04_30.
File
Use
instruction_dataset_main.jsonl
Full structured instruction upload (1680 rows)
instruction_dataset_smoke.jsonl
50-row smoke test, 5 languages x 5 formats x 2 rows
instruction_dataset_smoke_v2.jsonl
Same smoke set for the previous filename expected by notes
preference_pairs_scams.jsonl
Mirrored genuine-vs-scam… See the full description on the dataset page: https://huggingface.co/datasets/sahilmaniyar888/Indian_Climate_Disaster_data.Indian_Legal_NER_Datasetindian-cultural-datasetIndian_Legal_DatasetVakil_Indian_Legal_Datasetindian-law-datasetIndian_Law_Dataset_MinorProjectgg
indian-law-datasetindian-farmer-negotiation-data
🌾 Indian Farmer Mandi Negotiation Dataset
A high-quality, realistic training dataset for building AI systems that help Indian farmers negotiate better prices with traders at mandis (agricultural markets).
Dataset Details
Size: 5,000 examples
Language: Hindi / Hinglish (natural spoken style)
Coverage: 30 crops × 18 Indian states
Format: Input–Output pairs for supervised fine-tuning
Input Fields
Each example's input contains:
Field
Description
Example… See the full description on the dataset page: https://huggingface.co/datasets/StackOverflowed512/indian-farmer-negotiation-data.Indian-Law-AI-dataset
INDIAN LAW AI DATASET
This is the dataset used by the Indian Law AI system for legal text retrieval.
It consists of the following:-
acts.jsonl : This json file containing cleaned 892 Indian Acts.
judgements_new_m3.jsonl : This json file contains all the supreme court judgements where each case is divided into overlapping chunks to improve semantic retrieval.
acts_faiss_m3.faiss : This is a FAISS dataset that has the vector embeddings of the Indian Acts. It is… See the full description on the dataset page: https://huggingface.co/datasets/therealankit/Indian-Law-AI-dataset.indian_legal_dataset_qnaIndian_Laws_Structured_Legal_Dataset
📚 Indian Legal Acts Dataset (Structured Sections)
🧾 Overview
This dataset provides structured, machine-readable legal text from major Indian statutes, including:
Bharatiya Nyaya Sanhita, 2023 (BNS)
Code of Criminal Procedure, 1973 (CrPC)
Code of Civil Procedure, 1908 (CPC)
Indian Evidence Act, 1872 (IEA)
Negotiable Instruments Act, 1881 (NIA)
Motor Vehicles Act, 1988 (MVA)
Indian Divorce Act, 1869 (IDA)
Each entry represents a section or chunk of a section… See the full description on the dataset page: https://huggingface.co/datasets/dheerajpabolu/Indian_Laws_Structured_Legal_Dataset.chat-indian-revenue-datasetindian-law-datasetindian-law-datasetIndian-desktop-agent-training-datasetIndian_finance_datasetindian_jewelery_dataindian_doctors_dataset
