datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-india-law
Open India Law
Open, structured Indian primary law - plus the scrapers that build it.
Every judgment of the Supreme Court of India and all 25 High Courts, the decisions of 15
tribunals and regulators, and Central, State and Union Territory legislation down to the
individual section. Normalized to one schema, exclusively from official government sources.
Volume
Period
Court judgments
12,848,644
1950 to 2025
Tribunal and regulator matters
813,168
1985 to 2026… See the full description on the dataset page: https://huggingface.co/datasets/vaquill/open-india-law.india-index-options-1m
India Index & Options - 1-minute OHLC
1-minute OHLCV(+OI) bars for NSE/BSE index spot and option chains: NIFTY, BANKNIFTY, SENSEX (~2021-2026).
Powers the open-source TradeMarkk backtester (https://thetrademarkk.com).
Educational use only. Provided as-is, no warranty. Verify against official exchange data before relying on it.
Structure
index/{SYMBOL}.parquet - 1-min spot OHLC per index.
options/{SYMBOL}/{EXPIRY}.parquet - 1-min OHLC per option contract (with… See the full description on the dataset page: https://huggingface.co/datasets/thetrademarkk/india-index-options-1m.minuszero-indian-autonomous-driving-dataset-v2
INDUS-AD: Indian Dataset of Unstructured Urban Scenes for Autonomous Driving
Overview
INDUS-AD is the largest publicly released Indian autonomous-driving dataset for end-to-end autonomous-driving research. Its name expands to Indian Dataset of Unstructured Urban Scenes for Autonomous Driving.
This gated dataset is the decoded companion to the Minus Zero Indian Urban Autonomous Driving Dataset. It provides directly usable camera MP4s, normalized sensor tables… See the full description on the dataset page: https://huggingface.co/datasets/gagandeepreehal/minuszero-indian-autonomous-driving-dataset-v2.indian-market-historical-ohlcv
Indian Market Data (NSE/BSE)
Production-grade historical OHLCV dataset for Indian financial markets.
Auto-updated daily via GitHub Actions.
Dataset Summary
Metric
Value
Total Files
2670
Total Size
287.7 MB
Last Updated
2026-09-24 06:25 UTC
Update Frequency
Daily (weekdays)
Source
Yahoo Finance via yfinance
Asset Coverage
Asset Type
Symbols
Size
Stocks
2620
279.1 MB
Indices
17
3.1 MB
Etfs
17
2.0 MB
Commodities… See the full description on the dataset page: https://huggingface.co/datasets/vishnun0027/indian-market-historical-ohlcv.indian-markets
tejhq/indian-markets
End-of-day data for every NSE and BSE listed equity, built straight from the exchanges' official bhavcopy. Six parquet trees: raw prices, corporate actions, back-adjusted prices, symbol history, derived metrics, and a survivorship-bias-free liquidity universe. Refreshed every trading day at 20:00 IST by an open pipeline.
No broker, no auth, no scraping. The same data is served at api.tejhq.dev and mirrored on Cloudflare R2.
Coverage… See the full description on the dataset page: https://huggingface.co/datasets/tejhq/indian-markets.indian-stock-market-minute-data
🇮🇳 Indian Stock Market Data: Minute & Daily (2000 - 2026)
📌 Overview
This is a high-performance financial dataset containing the historical price history of 2,500+ NSE Stocks and Indices.
The dataset has been sharded and optimized for high-speed training. Instead of thousands of tiny files, it is grouped into large ~1.5GB Parquet shards, making it ideal for fast streaming with the Hugging Face datasets library.
📊 Dataset Stats
Total Rows: ~715 Million… See the full description on the dataset page: https://huggingface.co/datasets/xxparthparekhxx/indian-stock-market-minute-data.indian-case-laws
Indian Case Laws
Open Indian case-law data for AI, search, and legal research.
This dataset is part of the KanoonGPT Open Legal Data Initiative - an effort to make Indian legal data easier to access, trace, and build on for open-source research, legal tech, and production AI systems.
KanoonGPT is building structured Indian legal data and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.
Repository: KanoonGPT/indian-case-laws… See the full description on the dataset page: https://huggingface.co/datasets/KanoonGPT/indian-case-laws.tco-india
TCO India — car + fuel modules
Car makes, models, variants and dated on-road prices from CarDekho and CarWale (cross-source prices validated by variant descriptor + city), plus daily fuel prices. Previously tco-fuel (now folded as fuel/ module here).
Config
File
Rows describe
car_brands
car/brands.parquet
car brands (with ownership group + country)
car_models
car/models.parquet
car models
car_variants
car/variants.parquet
car variants (specs)
car_prices… See the full description on the dataset page: https://huggingface.co/datasets/bhavabhuthi/tco-india.indian-legal-documents
Indian Legal Documents
Open Indian statutory and legal-document data for AI, search, and legal research.
This dataset is part of the KanoonGPT Open Legal Data Initiative — an effort to make Indian legal data easier to access, structure, and build on for open-source research, legal tech, and production AI systems.
KanoonGPT is building structured Indian legal datasets and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.… See the full description on the dataset page: https://huggingface.co/datasets/KanoonGPT/indian-legal-documents.indiana-layoffs-warn-act-notices-daily
Indiana WARN Act layoff notices — every filing we hold since 2008, one CSV, rebuilt daily
1,182 Indiana WARN notices — every one this dataset holds, back to 2008 — free to download in full: no paywalled years, no login, no account · most recent notice filed 2026-09-02
· state source last checked 2026-09-23T12:27Z · official source: Indiana Department of Workforce Development — WARN notices.
Indiana employers must file a WARN Act notice with the state before a qualifying
mass… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/indiana-layoffs-warn-act-notices-daily.indian-court-decisions
Indian Court Decisions
A large-scale dataset of Indian court decisions with full text, metadata, and outcome labels covering the Supreme Court of India and 25 High Courts (1950–2026).
Dataset Summary
Config
Train
Validation
Test
Total
high_courts
11,682,776
1,459,319
1,457,934
14,600,029
supreme_court
40,044
4,990
5,019
50,053
Total
14,650,082
This is one of the largest publicly available legal NLP datasets, containing over 14.6 million… See the full description on the dataset page: https://huggingface.co/datasets/overthelex/indian-court-decisions.Indian_Financial_Newsindian-stock-market-minute-data
🇮🇳 Indian Stock Market Data: Minute & Daily (2000 - 2026)
📌 Overview
This is a high-performance financial dataset containing the historical price history of 2,500+ NSE Stocks and Indices.
The dataset has been sharded and optimized for high-speed training. Instead of thousands of tiny files, it is grouped into large ~1.5GB Parquet shards, making it ideal for fast streaming with the Hugging Face datasets library.
📊 Dataset Stats
Total Rows: ~715 Million… See the full description on the dataset page: https://huggingface.co/datasets/rahulkrraj/indian-stock-market-minute-data.indian-government-schemes-2025
Indian Government Schemes Dataset 2026
Dataset Description
The most comprehensive structured dataset of Indian central and state government schemes — 4,693 schemes across all ministries and states, with machine-readable eligibility fields.
Maintained by SmartDuke Technologies · Coimbatore, Tamil Nadu, India
This dataset powers SchemeFit — India's government scheme finder for citizens and businesses.
What Makes This Different
Most existing Indian… See the full description on the dataset page: https://huggingface.co/datasets/smartduketech/indian-government-schemes-2025.indian-court-decisions
Indian Court Decisions
A large-scale dataset of Indian court decisions with full text, metadata, and outcome labels covering the Supreme Court of India and 25 High Courts (1950–2026).
Dataset Summary
Config
Train
Validation
Test
Total
high_courts
11,682,776
1,459,319
1,457,934
14,600,029
supreme_court
40,044
4,990
5,019
50,053
Total
14,650,082
This is one of the largest publicly available legal NLP datasets, containing over 14.6 million… See the full description on the dataset page: https://huggingface.co/datasets/rtarun789/indian-court-decisions.one-year-of-r-indiaThis corpus contains the complete data for the activity of the subreddit /r/India from Sep 30, 2020 to Sep 30, 2021.india-upi-ecosystem-2018-2025
India UPI Ecosystem Dataset (2018-2025)
Dataset Summary
This dataset analyzes India's UPI transaction ecosystem by combining district-level app and usage data, official NPCI benchmark statistics, and RBI macroeconomic cash indicators.It is a merged and enriched analytics dataset designed for market concentration studies, geographic adoption analysis, forecasting, and cash displacement research.
Data Sources
Source
What it contains
Why it was used… See the full description on the dataset page: https://huggingface.co/datasets/prasad-gade05/india-upi-ecosystem-2018-2025.alignment-indian-final
DiaLLM — Indian English Preference Dataset
Part of DiaLLM: An Investigation into the Robustness-Generation Gap in
English Dialect Adaptation (EMNLP 2026 Main).
18,402 preference pairs for Indian English (en-IN), used for explicit-thread
DPO/GRPO/GSPO training targeting this variety.
Construction
Built from the UltraFeedback preference dataset (Cui et al., 2023):
the originally-preferred completion is transformed into a dialectal variant
using Multi-VALUE (Ziems… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/alignment-indian-final.india-index-options-1m
India Index & Options - 1-minute OHLC
1-minute OHLCV(+OI) bars for NSE/BSE index spot and option chains: NIFTY, BANKNIFTY, SENSEX (~2021-2026).
Powers the open-source TradeMarkk backtester (https://thetrademarkk.com).
Educational use only. Provided as-is, no warranty. Verify against official exchange data before relying on it.
Structure
index/{SYMBOL}.parquet - 1-min spot OHLC per index.
options/{SYMBOL}/{EXPIRY}.parquet - 1-min OHLC per option contract (with… See the full description on the dataset page: https://huggingface.co/datasets/codepyx23/india-index-options-1m.ipfs_india_laws_ir
India legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_india_laws (revision 051f15283bfc3f68bfb6926cf9f9899a8d3dff53) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of India prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was invented.… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_india_laws_ir.india-repo-rate-dataset
India RBI Policy Repo Rate and Monetary Policy Decision History
An independent, reproducible dataset of the Reserve Bank of India (RBI) policy
repo rate and its monetary-policy decision history. The canonical table is a
decision ledger: each row is one dated rate observation, including unchanged
policy decisions. Annual summaries, source metadata, contextual events, and
repository-provided regime intervals are separate configurations derived from
or accompanying that ledger.… See the full description on the dataset page: https://huggingface.co/datasets/ashwingopalsamy/india-repo-rate-dataset.indian-responsible-ai-benchmark
Indian Responsible AI Benchmark
A comprehensive benchmark for evaluating responsible AI behavior in Indian contexts — covering 212 adversarial and safety-critical prompts across 22 evaluation categories, 10 Indian language regions, and 8 Responsible AI dimensions.
Why This Benchmark?
Most AI safety benchmarks are US/Western-centric. Indian users face unique challenges:
Caste dynamics not captured by Western bias benchmarks
India/US context confusion (models… See the full description on the dataset page: https://huggingface.co/datasets/responsible-ai-labs/indian-responsible-ai-benchmark.IndianPersona-1M
IndianPersona-1M — Synthetic Indian Demographics & LLM Agent Personas
1,000,000 culturally-grounded synthetic Indian demographic profiles plus
250,000 ready-to-use LLM agent personas — generated entirely with the open-source
indian-fakedata library
(PyPI · npm).
This dataset is 100% synthetic. Every row carries synthetic = true.
All identifiers (Aadhaar, PAN, voter ID, phone, email) are fabricated and exist in no
government or commercial database. No real individual is… See the full description on the dataset page: https://huggingface.co/datasets/Abhay557/IndianPersona-1M.census-2011
India Census 2011 — District Level
Clean, versioned district-level data from India's 2011 Census.
640 districts. 29 columns. Zero missing values. Ready for pandas.
Quick Start
import pandas as pd
from huggingface_hub import hf_hub_download
path = hf_hub_download(
repo_id="indiaset/census-2011",
filename="census_2011_districts_final.parquet",
repo_type="dataset"
)
df = pd.read_parquet(path)
print(df.shape) # (640, 29)
What's In This… See the full description on the dataset page: https://huggingface.co/datasets/indiaset/census-2011.india-solar-benchmark-dataset
India Solar Benchmark Dataset
A large-scale benchmark dataset for solar irradiance forecasting and renewable energy research built from NASA POWER meteorological observations across 50 major Indian cities.
The benchmark contains 10 years of hourly observations (2016–2025) and is distributed as two complementary datasets:
india_multicity_raw.parquet – cleaned and standardized observations after preprocessing, intended for custom feature engineering and research.… See the full description on the dataset page: https://huggingface.co/datasets/Narendersingh007/india-solar-benchmark-dataset.dexfluence-indian-creator-index
Dexfluence Indian Creator Index
Verified Indian influencer dataset across Instagram, YouTube, and TikTok with engagement rates, follower tier, niche classification, and authenticity scores.
Dataset summary
141,000+ verified Indian creators indexed across Instagram, YouTube, and TikTok
Top 5,000 by follower count included in this Hugging Face mirror (CC-BY 4.0)
Each record includes: handle, name, platform, niche, follower count, engagement rate, country… See the full description on the dataset page: https://huggingface.co/datasets/Shikha180224/dexfluence-indian-creator-index.blue-block-so101-50ep-swapedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/IndiaTechTeamSL2/blue-block-so101-50ep-swaped.indian-stock-market-minute-data
🇮🇳 Indian Stock Market Data: Minute & Daily (2000 - 2026)
📌 Overview
This is a high-performance financial dataset containing the historical price history of 2,500+ NSE Stocks and Indices.
The dataset has been sharded and optimized for high-speed training. Instead of thousands of tiny files, it is grouped into large ~1.5GB Parquet shards, making it ideal for fast streaming with the Hugging Face datasets library.
📊 Dataset Stats
Total Rows: ~715… See the full description on the dataset page: https://huggingface.co/datasets/GalacticWanderer/indian-stock-market-minute-data.2019_Major_Indian_Airlines_Datapima-indians-diabetes-database
Pima Indians Diabetes Dataset Split
This directory contains split datasets of Pima Indians Diabetes Database.
For each splits, we have
Mock data: The mock data is a smaller dataset (10 rows for both train and test) that is used to test the model and data processing code.
Private data: Each private data contains 123-125 rows for training, and 32-33 rows for testing.
