datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
minuszero-indian-autonomous-driving-dataset-v2
INDUS-AD: Indian Dataset of Unstructured Urban Scenes for Autonomous Driving
Overview
INDUS-AD is the largest publicly released Indian autonomous-driving dataset for end-to-end autonomous-driving research. Its name expands to Indian Dataset of Unstructured Urban Scenes for Autonomous Driving.
This gated dataset is the decoded companion to the Minus Zero Indian Urban Autonomous Driving Dataset. It provides directly usable camera MP4s, normalized sensor tables… See the full description on the dataset page: https://huggingface.co/datasets/gagandeepreehal/minuszero-indian-autonomous-driving-dataset-v2.indian-stock-market-minute-data
🇮🇳 Indian Stock Market Data: Minute & Daily (2000 - 2026)
📌 Overview
This is a high-performance financial dataset containing the historical price history of 2,500+ NSE Stocks and Indices.
The dataset has been sharded and optimized for high-speed training. Instead of thousands of tiny files, it is grouped into large ~1.5GB Parquet shards, making it ideal for fast streaming with the Hugging Face datasets library.
📊 Dataset Stats
Total Rows: ~715 Million… See the full description on the dataset page: https://huggingface.co/datasets/xxparthparekhxx/indian-stock-market-minute-data.indian-legal-supervised-fine-tuning-data
🇮🇳 LegalBrain Indic Legal Corpus
A large-scale multilingual Indian legal dataset curated to support research in:
Domain-specific LLM training
Legal question answering
Policy reasoning & case retrieval
Agentic systems for legal workflow automation
This dataset contains text drawn from publicly available legal sources across multiple Indian languages, including:
English, Hindi, Marathi, Bengali, Kannada, Tamil, Telugu, Odia, and others.
The corpus is structured and processed to be… See the full description on the dataset page: https://huggingface.co/datasets/Prarabdha/indian-legal-supervised-fine-tuning-data.minuszero-indian-autonomous-driving-dataset
Minus Zero Indian Urban Autonomous Driving Dataset
Overview
This dataset provides original multicamera autonomous-driving recordings in MCAP format. It is designed for non-commercial research on surround-view perception, temporal and cross-camera synchronization, H.265 video pipelines, localization, GNSS/pose integration, and robotics data tooling.
Recordings include camera and GNSS/pose streams, with machine-state telemetry present in a small subset. Camera… See the full description on the dataset page: https://huggingface.co/datasets/gagandeepreehal/minuszero-indian-autonomous-driving-dataset.indian-law-datasetindian-stock-market-minute-data
🇮🇳 Indian Stock Market Data: Minute & Daily (2000 - 2026)
📌 Overview
This is a high-performance financial dataset containing the historical price history of 2,500+ NSE Stocks and Indices.
The dataset has been sharded and optimized for high-speed training. Instead of thousands of tiny files, it is grouped into large ~1.5GB Parquet shards, making it ideal for fast streaming with the Hugging Face datasets library.
📊 Dataset Stats
Total Rows: ~715 Million… See the full description on the dataset page: https://huggingface.co/datasets/rahulkrraj/indian-stock-market-minute-data.indian_cultural_raw_dataset
Indian Cultural Dataset
This dataset contains various Indian cultural elements including:
Cultural Elements
Folks and Regional Stories
Historical Events
Mythology
Regional Elements
Value Systems & Teachings
Dataset Structure
The dataset is organized into the following directories:
Cultural Elements/
folks and regional stories/
Historical events/
mythology/
Regional element/
Value Systems & Teachings/
Content
The dataset includes PDF and text files… See the full description on the dataset page: https://huggingface.co/datasets/ombhojane/indian_cultural_raw_dataset.Indian-Multilingual-Bias-Dataset
Indian Multilingual Bias Dataset
Dataset Description
The Indian Multilingual Bias Dataset is a comprehensive collection designed to evaluate and measure social biases in Large Language Models (LLMs) across three major Indian languages: English, Bengali (বাংলা), and Hindi (हिंदी). This dataset is based on the original Indian-BhED dataset and focuses on four critical dimensions of bias prevalent in Indian society.
Key Features
🌐 Multilingual:… See the full description on the dataset page: https://huggingface.co/datasets/Debk/Indian-Multilingual-Bias-Dataset.ILID_Indian_Language_Identification_Dataset
ILID: Native Script Language Identification for Indian Languages
Paper | Code | Project Page
🗣 ILID: Indian Language Identification Dataset (23 Languages)Authors: Yash Ingle, Dr. Pruthwik MishraInstitute: Sardar Vallabhbhai National Institute of Technology (SVNIT), Surat, India
📄 Dataset Description
The ILID (Indian Language Identification Dataset) benchmark contains 250,000sentences from English and 22 official Indian languages, designed for training and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/yash-ingle/ILID_Indian_Language_Identification_Dataset.Indian-Accent-Datasetindian-pharma-dataset-2026-augast
PharmaLens: 200k Medicine Catalog (Salts, Prices, Interactions and Reviews)
I spent weeks compiling and cleaning this retrieval database for a project. Instead of letting 300MB+ of structured pharmaceutical data sit idle on my hard drive, I am open-sourcing it. Use it for your RAG pipelines, chatbots, pricing tools, or whatever else you are building.
Overview
Finding clean, structured pharmaceutical datasets with commercial brand names, active salt compositions… See the full description on the dataset page: https://huggingface.co/datasets/sinhal/indian-pharma-dataset-2026-augast.Indian-legal-data-v3
Indian Legal Dataset V3
Overview
Indian Legal Dataset V3 is a large-scale instruction-tuning dataset focused on Indian law, constitutional law, criminal law, legal reasoning, legal drafting, and real-world legal assistance.
Compared to V2, this version expands the dataset with:
legal drafting instruction pairs,
hypothetical legal scenarios,
detailed IPC-focused data,
practical real-world legal instructions,
concise legal QA pairs.
After integrating the new data sources… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/Indian-legal-data-v3.indian-stock-market-minute-data
🇮🇳 Indian Stock Market Data: Minute & Daily (2000 - 2026)
📌 Overview
This is a high-performance financial dataset containing the historical price history of 2,500+ NSE Stocks and Indices.
The dataset has been sharded and optimized for high-speed training. Instead of thousands of tiny files, it is grouped into large ~1.5GB Parquet shards, making it ideal for fast streaming with the Hugging Face datasets library.
📊 Dataset Stats
Total Rows: ~715… See the full description on the dataset page: https://huggingface.co/datasets/GalacticWanderer/indian-stock-market-minute-data.indian-legal-opposing-counsel-dataset
⚖️ Indian Legal Opposing Counsel Dataset
A combined, preprocessed dataset of 26,326 examples for training an Indian legal opposing counsel AI model. Ready-to-use in ChatML format for SFT training.
📊 Dataset Stats
Split
Rows
Size
Train
25,009
65 MB
Test
1,317
3.5 MB
Total
26,326
69 MB
📦 Sources
Source Dataset
Rows
Content
viber1/indian-law-dataset
24,607
Writs, PIL, civil procedure, constitutional law, IPC… See the full description on the dataset page: https://huggingface.co/datasets/pkheria7/indian-legal-opposing-counsel-dataset.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/ysangam/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.flipkart-data-indianAnna-Data-Indian-Culinary-Dataset
Anna-Data: Multilingual Indian Culinary Dataset v0.0.1
A comprehensive, validated, multilingual dataset of 100+ Indian ingredients across 7 categories. Each ingredient is annotated with botanical names, nutritional information, and culinary attributes. This dataset serves as a foundational resource for AI applications in Indian cuisine, nutrition analysis, and cultural studies.
🚀 Dataset Highlights
Multilingual Support: 11 languages (English + 10 Indian languages).… See the full description on the dataset page: https://huggingface.co/datasets/nandhar/Anna-Data-Indian-Culinary-Dataset.Indian-Legal-SFT-Dataset
Vidhaan: High-Density Indian Legal Instruction Dataset
Vidhaan is a comprehensive, high-precision instruction-tuning dataset containing 20,690 QA pairs derived from 113 Central Acts of India. It was built specifically to solve the "context-splitting" problem found in standard legal RAG datasets.
🛠 Dataset Structure & Format
Primary File: vidhaan_training_v1.jsonl
Format: JSON Lines (JSONL)
Schema: - instruction: (String) A precise legal query.
context: (String) The… See the full description on the dataset page: https://huggingface.co/datasets/SharathReddy/Indian-Legal-SFT-Dataset.indian-pharma-dataindian-legal-records
LH2 Data — Indian Legal Records & Judgments Corpus
The most comprehensive structured Indian legal records corpus available for AI training — 267M+ case records spanning the full judicial hierarchy, paired with a pre-computed AI enrichment layer across 21M+ court orders.
Dataset Summary
This corpus provides structured, indexed, and partially labelled legal records from the Indian judicial system at a scale that has no public equivalent. It covers the Supreme Court of… See the full description on the dataset page: https://huggingface.co/datasets/LH2-data-labs/indian-legal-records.Indian-legal-data-v2
Legal Instruction Dataset (v2)
📌 Overview
This dataset contains high-quality instruction–response pairs derived from Indian legal texts, primarily focusing on statutory interpretation and structured legal explanations.
Version 2 represents a significant scale and quality upgrade over v1:
v1: 33,077 samples
v2: 171,640 samples
The dataset is designed specifically for instruction tuning of language models, emphasizing clarity, structure, and legal reasoning patterns.… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/Indian-legal-data-v2.indian_ipo_prospectus_data_with_pageno
Dataset Card for Dataset Name
Dataset Summary
Prospectus text mining is very important for the investor community to identify major risks.
factors and evaluate the use of the amount to be raised during an IPO. For this dataset author
downloaded 100 prospectuses from the Indian Market Regulator website. The dataset contains the URL and OCR text for 100 prospectuses.
Further, the author released a Roberta LM and sentence transformer for usage.
This dataset Contains Page… See the full description on the dataset page: https://huggingface.co/datasets/scholarly360/indian_ipo_prospectus_data_with_pageno.indian-recipe-datasetIndian_Names_with_Gender_Dataset
🇮🇳 Indian Names & Gender Dataset (Balanced)
Dataset Summary
This dataset contains 42,000 samples designed for training models to identify Indian names and classify their gender. It is perfectly balanced across three categories, making it ideal for training robust classifiers that can distinguish between real names and random text.
Dataset Structure
Column
Type
Description
Name
String
The text string (Name or Random Word).
Label
Integer
Class ID… See the full description on the dataset page: https://huggingface.co/datasets/shisha-07/Indian_Names_with_Gender_Dataset.Indian-legal-data-v1
Legal Instruction Dataset
📌 Overview
This dataset contains instruction–response pairs derived from sections of the Indian Acts.
The dataset is designed for instruction tuning of language models, with a focus on:
structured legal explanations
bullet-point formatting
long-form responses
🧠 Dataset Description
Task Type: Instruction Tuning / Legal QA
Domain: Indian Law
Language: English
Format: JSONL
🔥 What makes this good (not generic fluff)… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/Indian-legal-data-v1.indian_voices_dataset_allIndian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/DatasetNewUser/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.Indian_Traffic_VQA_Dataset🧭 Overview
Indian Traffic VQA is a real-world Visual Question Answering (VQA) dataset focusing on Indian road traffic signboards.
The dataset is designed for training and evaluating Vision-Language Models (VLMs) and VQA systems in the traffic and transportation domain.
This dataset bridges a gap between real-world Indian traffic conditions and machine understanding — ideal for research in autonomous driving, smart city AI, and traffic sign recognition under natural environments.
📦 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/chandrabhuma/Indian_Traffic_VQA_Dataset.indian_ipo_prospectus_data
Dataset Card for Dataset Name
Dataset Summary
Prospectus text mining is very important for the investor community to identify major risks.
factors and evaluate the use of the amount to be raised during an IPO. For this dataset author
downloaded 100 prospectuses from the Indian Market Regulator website. The dataset contains the URL and OCR text for 100 prospectuses.
Further, the author released a Roberta LM and sentence transformer for usage.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/scholarly360/indian_ipo_prospectus_data.indian-tts-dataset
Indian TTS Dataset
A curated Text-to-Speech training dataset with high-quality audio clips,
transcriptions, and emotion labels for Indian English (en-IN) and Hindi (hi-IN).
Dataset Summary
Metric
Value
Total clips
125
English (en-IN)
96 clips
Hindi (hi-IN)
29 clips
Total duration
28.8 minutes
Sample rate
22050 Hz
Format
WAV (PCM 16-bit, mono)
Emotion Distribution
Emotion
Count
neutral
92
narrative
10
excited… See the full description on the dataset page: https://huggingface.co/datasets/champTUSHARg007/indian-tts-dataset.
