datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IndicVoices
IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages
Updates
[23 December 2025] We now have 11,200 hours of transcribed data! 🎉
Overview
INDICVOICES is a dataset of natural and spontaneous speech containing a total of 23.7K hours of read (8%), extempore (76%) and conversational (15%) audio from 51K speakers covering 400+ Indian districts and 22 languages. Of these 23.7K hours, 11.2K hours have… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicVoices.open-india-law
Open India Law
Open, structured Indian primary law - plus the scrapers that build it.
Every judgment of the Supreme Court of India and all 25 High Courts, the decisions of 15
tribunals and regulators, and Central, State and Union Territory legislation down to the
individual section. Normalized to one schema, exclusively from official government sources.
Volume
Period
Court judgments
12,848,644
1950 to 2025
Tribunal and regulator matters
813,168
1985 to 2026… See the full description on the dataset page: https://huggingface.co/datasets/vaquill/open-india-law.dojo_fin_indicators
Languages: 简体中文 · English
dojo_fin_indicators — Financial Metrics
Overview
Multi-period financial statement derivatives per symbol: income statement, balance sheet, cash flow, and industry-specific metrics (banks, insurers, brokers, etc.). Supports quarterly, cumulative, and other report_type values.
Files
File
Description
data.parquet
Full financial metrics (wide table, 100+ columns)
Key Fields (common)
Field… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_fin_indicators.Indian-Laws
Dataset Card for Indian Laws
This is a comprehensive collection of primary legal documents pertinent to the Indian legal system.
It is designed to serve as a foundational resource for supervised fine-tuning (SFT) to make language models, particularly those focused on legal applications tailored for Indian law.
india-index-options-1m
India Index & Options - 1-minute OHLC
1-minute OHLCV(+OI) bars for NSE/BSE index spot and option chains: NIFTY, BANKNIFTY, SENSEX (~2021-2026).
Powers the open-source TradeMarkk backtester (https://thetrademarkk.com).
Educational use only. Provided as-is, no warranty. Verify against official exchange data before relying on it.
Structure
index/{SYMBOL}.parquet - 1-min spot OHLC per index.
options/{SYMBOL}/{EXPIRY}.parquet - 1-min OHLC per option contract (with… See the full description on the dataset page: https://huggingface.co/datasets/thetrademarkk/india-index-options-1m.thai-commoncrawl-index
Thai Common Crawl Index (2019–2026)
An index of every page Common Crawl detected as Thai across 70 monthly crawls, from
January 2019 (CC-MAIN-2019-04) to August 2026 (CC-MAIN-2026-30).
932,874,727 page captures · 450,971,497 unique URLs · 6,997,185 hosts · 6,674,969 domains
Each row records where the page lives inside Common Crawl's WARC archives — file name,
byte offset, and record length — so you can fetch exactly the pages you want with HTTP
range requests, without scanning… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-commoncrawl-index.indicvoices_r
IndicVoices-R: Multilingual, Multi-Speaker Speech Corpus for Indian TTS
Dataset Summary
IndicVoices-R (IV-R) is the largest multilingual Indian text-to-speech (TTS) dataset derived from an automatic speech recognition (ASR) dataset. It contains 1,704 hours of high-quality speech from 10,496 speakers across 22 Indian languages. This dataset is designed to enhance the development of robust Indian TTS models by providing diverse speaker demographics, natural… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/indicvoices_r.host-index-testing-v2
Common Crawl Host Index v2
GitHub: https://github.com/commoncrawl/cc-host-index
Each crawl, we generate a Host Index, which aggregates information about each web hosted visited during the crawl. The
information is aggregated from the Common Crawl columnar index,
web graph, and raw crawler logs.
Quickstart
The dataset is Hive-partitioned on crawl (data/crawl=CC-MAIN-2025-18/*.parquet). Open the whole
dataset once, then filter with WHERE crawl = '...': because… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/host-index-testing-v2.IndustrialDetectionStaticCamerasThe IndustrialDetectionStaticCameras dataset has been collected in order to validate the methodology presented in the paper entitled A few-shot learning methodology for improving safety in industrial scenarios through universal self-supervised visual features and dense optical flow. This dataset is divided into five main folders named videoY, where Y=1,2,3,4,5. Each videoY folder contains the following:
The video of the scene in .mp4 format: videoY.mp4
A folder with the images of each frame… See the full description on the dataset page: https://huggingface.co/datasets/jjldo21/IndustrialDetectionStaticCameras.minuszero-indian-autonomous-driving-dataset-v2
INDUS-AD: Indian Dataset of Unstructured Urban Scenes for Autonomous Driving
Overview
INDUS-AD is the largest publicly released Indian autonomous-driving dataset for end-to-end autonomous-driving research. Its name expands to Indian Dataset of Unstructured Urban Scenes for Autonomous Driving.
This gated dataset is the decoded companion to the Minus Zero Indian Urban Autonomous Driving Dataset. It provides directly usable camera MP4s, normalized sensor tables… See the full description on the dataset page: https://huggingface.co/datasets/gagandeepreehal/minuszero-indian-autonomous-driving-dataset-v2.IndicSynth
IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset
A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research
🏆 Outstanding Paper Award, ACL 2025
🧠 Overview
IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/vdivyasharma/IndicSynth.Indian-Laws
Dataset Card for Indian Laws
This is a comprehensive collection of primary legal documents pertinent to the Indian legal system.
It is designed to serve as a foundational resource for supervised fine-tuning (SFT) to make language models, particularly those focused on legal applications tailored for Indian law.
NCERT-Parallel-Dataset-Indicindic_glue
Dataset Card for "indic_glue"
Dataset Summary
IndicGLUE is a natural language understanding benchmark for Indian languages. It contains a wide
variety of tasks and covers 11 major Indian languages - as, bn, gu, hi, kn, ml, mr, or, pa, ta, te.
The Winograd Schema Challenge (Levesque et al., 2011) is a reading comprehension task
in which a system must read a sentence with a pronoun and select the referent of that pronoun from
a list of choices. The examples are manually… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/indic_glue.indoor-safety-hazard-detection-and-work-zone-monitoring
Indoor Safety Hazard Detection & Work-Zone Monitoring
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by… See the full description on the dataset page: https://huggingface.co/datasets/physicl/indoor-safety-hazard-detection-and-work-zone-monitoring.IndustryCorpus2_tourism_geography
IndustryCorpus2: Travel & Geography
This repository contains the IndustryCorpus2: Travel & Geography domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao},
year… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_tourism_geography.vidore_v3_industrialViDoRe V3 : Industrial reports
This dataset, Industrial reports, is a corpus of technical documents on military aircrafts (fueling, mechanics...), intended for complex-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark.
About ViDoRe v3
ViDoRe V3 is our latest benchmark for RAG evaluation on visually-rich documents from real-world applications. It features 10 datasets with, in total, 26,000 pages and 3099 queries, translated into 6… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_industrial.indic-dialect-asr
Indic Dialect ASR Dataset
A multilingual ASR dataset covering 30 Indic dialect/languages with 2.8M+ samples.
Usage
from datasets import load_dataset
# Load a specific language
ds = load_dataset("grushaaaaa/indic-dialect-asr", "assamese", split="train")
Features
audio: 16kHz WAV audio
sentence: Transcription text
language: Language name
source: Source dataset
polymarket-search-indexInddataIndicSynth
IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset
A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research
🏆 Outstanding Paper Award, ACL 2025
🧠 Overview
IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/ksmashhero/IndicSynth.Indices-Daily-Price
Indices Daily Price
This dataset includes daily price data for various indices.
815,441 rows over 113 symbols, 8 columns, covering 1927-12-30 to 2026-08-03. Refreshed monthly.
Strategies Built on This Data
1,324 papers in the Papers With Backtest catalogue declare this dataset as an input. 1,226 of them have been coded and run over their own full history. The median replicated Sharpe ratio is +0.49, and 62% clear a t-statistic of 1.96 on their own sample, against… See the full description on the dataset page: https://huggingface.co/datasets/paperswithbacktest/Indices-Daily-Price.umi-robots
umi-robots
Part of umi, an open web crawl published as Parquet. Before you use any of this, read the exclusion list at open-index/umi-meta and filter the rows it names. Published files are never rewritten, so the exclusion list is how a takedown reaches you, and applying it is a condition of using the data rather than a suggestion.
One row per robots.txt fetch: the host, when we asked, what the origin answered, the raw text if it served one, and the summary our parser read out… See the full description on the dataset page: https://huggingface.co/datasets/open-index/umi-robots.IndicGenBenchFloresBitextMining
IndicGenBenchFloresBitextMining
An MTEB dataset
Massive Text Embedding Benchmark
Flores-IN dataset is an extension of Flores dataset released as a part of the IndicGenBench by Google
Task category
t2t
Domains
Web, News, Written
Reference
https://github.com/google-research-datasets/indic-gen-bench/
Source datasets:
google/IndicGenBench_flores_in
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:… See the full description on the dataset page: https://huggingface.co/datasets/mteb/IndicGenBenchFloresBitextMining.Equal_Industry_day_10_pqtIndian-Laws
Dataset Card for Indian Laws
This is a comprehensive collection of primary legal documents pertinent to the Indian legal system.
It is designed to serve as a foundational resource for supervised fine-tuning (SFT) to make language models, particularly those focused on legal applications tailored for Indian law.
indic-tts-966h
Indic-TTS-966h
Six-language Indian TTS corpus: ~966 hours of paired speech and text, 24 kHz mono WAV
clips with sentence-level transcripts in native scripts (natural English code-switching
preserved).
Subset
Clips
Hours
bengali
18,343
94.9
malayalam
30,548
192.5
marathi
34,327
213.4
punjabi
28,083
161.8
tamil
26,817
171.1
telugu
21,923
132.8
Columns: audio (24 kHz mono), file_name, transcript. One config per language:
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/psk/indic-tts-966h.IndustryCorpus2_medicine_health_psychology_traditional_chinese_medicine
IndustryCorpus2: Health & Medicine
This repository contains the IndustryCorpus2: Health & Medicine domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao},
year… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_medicine_health_psychology_traditional_chinese_medicine.Inddataindic-align
IndicAlign
A diverse collection of Instruction and Toxic alignment datasets for 14 Indic Languages. The collection comprises of:
IndicAlign - Instruct
Indic-ShareLlama
Dolly-T
OpenAssistant-T
WikiHow
IndoWordNet
Anudesh
Wiki-Conv
Wiki-Chat
IndicAlign - Toxic
HHRLHF-T
Toxic-Matrix
We use IndicTrans2 (Gala et al., 2023) for the translation of the datasets.
We recommend the readers to check out our paper on Arxiv for detailed information on the curation process of these… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/indic-align.
