datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IndianBailJudgments-1200
⚖️ IndianBailJudgments-1200: Annotated Indian Bail Order Dataset (1975–2025)
IndianBailJudgments-1200 is a high-quality, structured dataset comprising 1,200 annotated Indian bail-related court orders spanning five decades (1975–2025). It captures granular legal information across 78 courts and 28 regions, including crime types, IPC sections invoked, judge names, legal issues, bail outcomes, and bias indicators.
Designed for use in legal NLP, fairness analysis, and judicial research… See the full description on the dataset page: https://huggingface.co/datasets/SnehaDeshmukh/IndianBailJudgments-1200.India-Stock-Symbols-and-Metadata
India Stock Symbols & Company Metadata
This dataset contains stock symbols and basic company metadata for all listed companies in India.It is updated weekly if new changes are there.
📊 Dataset Contents
The dataset is provided as a CSV file with the following columns:
Column
Description
name
Full company name
ticker
Stock ticker symbol (e.g., AAPL, MSFT)
market
The exchange/market where the stock is listed
sector
The primary business sector of the… See the full description on the dataset page: https://huggingface.co/datasets/kjhq/India-Stock-Symbols-and-Metadata.indiana-layoffs-warn-act-notices-daily
Indiana WARN Act layoff notices — every filing we hold since 2008, one CSV, rebuilt daily
1,182 Indiana WARN notices — every one this dataset holds, back to 2008 — free to download in full: no paywalled years, no login, no account · most recent notice filed 2026-09-02
· state source last checked 2026-09-23T12:27Z · official source: Indiana Department of Workforce Development — WARN notices.
Indiana employers must file a WARN Act notice with the state before a qualifying
mass… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/indiana-layoffs-warn-act-notices-daily.nse-india-security-master
TickerTruth — NSE India Security Master (Explorer)
A clean, normalized reference table of 2,389 NSE-listed equities — ISIN mappings, listing dates, company names, and active/delisted status — built from the TickerTruth India reference-data pipeline.
Why this dataset exists
India equity data is notoriously messy. NSE symbols change (renames, mergers, delistings), ISINs get reissued, and raw bhavcopy files carry no historical context. TickerTruth's pipeline… See the full description on the dataset page: https://huggingface.co/datasets/tickertruthorg/nse-india-security-master.indian-premier-leagueIndian Premier League Dataset
This dataset contains info on all of the IPL(Indian Premier League) cricket matches.
Ball-by-Ball level info and scorecard info to be added soon.
The dataset was scraped in July-2022.
Mantainers:
Somya Gautam
Kondrolla Dinesh Reddy
Keshaw Soni
indian-government-schemes-2025
Indian Government Schemes Dataset 2026
Dataset Description
The most comprehensive structured dataset of Indian central and state government schemes — 4,693 schemes across all ministries and states, with machine-readable eligibility fields.
Maintained by SmartDuke Technologies · Coimbatore, Tamil Nadu, India
This dataset powers SchemeFit — India's government scheme finder for citizens and businesses.
What Makes This Different
Most existing Indian… See the full description on the dataset page: https://huggingface.co/datasets/smartduketech/indian-government-schemes-2025.nse-india-security-master
TickerTruth — NSE India Security Master (Explorer)
A clean, normalized reference table of 2,389 NSE-listed equities — ISIN mappings, listing dates, company names, and active/delisted status — built from the TickerTruth India reference-data pipeline.
Why this dataset exists
India equity data is notoriously messy. NSE symbols change (renames, mergers, delistings), ISINs get reissued, and raw bhavcopy files carry no historical context. TickerTruth's pipeline… See the full description on the dataset page: https://huggingface.co/datasets/tickertruth/nse-india-security-master.synthetic-indian-logical-reasoning-CoTyes
indian-legalIndian_Financial_News
Dataset Card for Dataset Name
The FinancialNewsSentiment_26000 dataset comprises 26,000 rows of financial news articles related to the Indian market. It features four columns: URL, Content (scrapped content), Summary (generated using the T5-base model), and Sentiment Analysis (gathered using the GPT add-on for Google Sheets). The dataset is designed for sentiment analysis tasks, providing a comprehensive view of sentiments expressed in financial news.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/kdave/Indian_Financial_News.llama2_indian_law_v1IndianDomesticAirlineDatasetIndian Domestic Airline Flights (2018–2025)
This dataset contains Indian domestic airline [103 - Airports] data covering the years 2018 to 2025.
It includes flight numbers, airlines, and other relevant attributes for domestic routes within India.
This dataset can be used for:
Flight delay prediction
Airline trend analysis
Route popularity insights
Column Descriptions
Column Name
Descriptions
Airline
Name of the airline operating the flight (e.g., IndiGo, Air India).
FlightNumber… See the full description on the dataset page: https://huggingface.co/datasets/Kabil007/IndianDomesticAirlineDataset.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/ysangam/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.IndianLegal-QA
IndianLegal-QA
A question-and-answer dataset derived from Indian legal and government documents,
covering the Constitution of India, the Indian Penal Code, criminal and civil
procedural law, customs and tariff classifications, and numerous central and
state acts. The dataset is suitable for building, fine-tuning, and evaluating
retrieval and question-answering systems over Indian legal text.
This dataset is also hosted on GitHub at Sakib-Dalal/IndianLegal-QA.… See the full description on the dataset page: https://huggingface.co/datasets/Sakib-Dalal/IndianLegal-QA.scam-hum-india
Scam/Spam India Dataset
A dataset of labeled text messages for scam/spam detection, focused on Indian scam patterns
(telecom promotions, lottery fraud, Aadhaar/KYC phishing, UPI/bank fraud, fake job offers, OTP-theft attempts).
Dataset Structure
Column
Type
Description
text
string
The message content
label
string
ham (legitimate), spam/scam
Total rows: 2272
Label distribution: {'ham': 1377, 'spam': 895}
Source
Merged from two… See the full description on the dataset page: https://huggingface.co/datasets/anmolshrivastav/scam-hum-india.india-r-1-year-postsindian-passports
Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes.
Introduction
The Synthetic India Passports Dataset assembles more than 1,000 AI-generated passport images intended for training OCR and computer vision models on identity documents. Because every record is fully synthetic — with no… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/indian-passports.indian-transaction-categorization-synthetic
Synthetic Indian Bank Transaction Narrations
810 synthetic (text, category) pairs mimicking Indian bank/credit-card statement narrations —
built to train the Sumeetgpt/indian-transaction-categorizer
SetFit model.
Why this exists
While building a personal finance app, we searched for a public dataset pairing real Indian
transaction narration formats (UPI, NEFT, IMPS, ACH) with spending-category labels, and found
none: datasets with real-looking Indian narration… See the full description on the dataset page: https://huggingface.co/datasets/Sumeetgpt/indian-transaction-categorization-synthetic.amazon_india_productsIndian-Constitution
Indian Constitution Dataset
The dataset can be used for text classification, text generation and text2text generation
Indian-Multilingual-Bias-Dataset
Indian Multilingual Bias Dataset
Dataset Description
The Indian Multilingual Bias Dataset is a comprehensive collection designed to evaluate and measure social biases in Large Language Models (LLMs) across three major Indian languages: English, Bengali (বাংলা), and Hindi (हिंदी). This dataset is based on the original Indian-BhED dataset and focuses on four critical dimensions of bias prevalent in Indian society.
Key Features
🌐 Multilingual:… See the full description on the dataset page: https://huggingface.co/datasets/Debk/Indian-Multilingual-Bias-Dataset.indian-pharma-dataindian-stock-hourly-2017-2021
Indian Stock Market Hourly Data (2017-2021)
Hourly OHLCV data for major Indian stock market indices and their futures contracts from the National Stock Exchange (NSE).
Overview
Metric
Value
Time Period
May 2017 - December 2021
Frequency
Hourly (8 candles/day)
Instruments
5 (2 futures + 3 indices)
Total Data Points
40,445 hourly bars
Coverage
~85-90% of trading hours
Splits
Split
Instrument
Description
Rows
nifty_fut
NIFTY… See the full description on the dataset page: https://huggingface.co/datasets/calender/indian-stock-hourly-2017-2021.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/DatasetNewUser/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.indian-recipe-datasetabhinand05_crop-production-in-india
Crop Production in India
Can you predict crop production in India?
Dataset Info
Source: Kaggle
Original Size: 1.96 MB
Kaggle Downloads: 17,273
Files: 1
Files
crop_production.csv
Mirrored from Kaggle
flipkart-data-indianLaws_and_Constitution_of_Indiaindian_names
indian_names
Unique customer names (lowercased) extracted from call/loan record databases
(xPertVoice Aug/Sep, both "PQ" and "PQ2" variants), each paired with the
language(s) detected from s3_key.
customer_names_with_language.csv — two columns: name, language.
One row per unique name. If a name was seen under multiple languages,
all of them are listed in one comma-separated field (e.g. "english, tamil").
Language is detected by scanning s3_key for a known Indian-language… See the full description on the dataset page: https://huggingface.co/datasets/MeghanaKap/indian_names.IndianDomesticAirlineDatasetIndian Domestic Airline Flights (2018–2025)
This dataset contains Indian domestic airline [103 - Airports] data covering the years 2018 to 2025.
It includes flight numbers, airlines, and other relevant attributes for domestic routes within India.
This dataset can be used for:
Flight delay prediction
Airline trend analysis
Route popularity insights
Column Descriptions
Column Name
Descriptions
Airline
Name of the airline operating the flight (e.g., IndiGo, Air India).
FlightNumber… See the full description on the dataset page: https://huggingface.co/datasets/Leo4664/IndianDomesticAirlineDataset.
