datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IndicVoices
IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages
Updates
[23 December 2025] We now have 11,200 hours of transcribed data! 🎉
Overview
INDICVOICES is a dataset of natural and spontaneous speech containing a total of 23.7K hours of read (8%), extempore (76%) and conversational (15%) audio from 51K speakers covering 400+ Indian districts and 22 languages. Of these 23.7K hours, 11.2K hours have… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicVoices.dojo_fin_indicators
Languages: 简体中文 · English
dojo_fin_indicators — Financial Metrics
Overview
Multi-period financial statement derivatives per symbol: income statement, balance sheet, cash flow, and industry-specific metrics (banks, insurers, brokers, etc.). Supports quarterly, cumulative, and other report_type values.
Files
File
Description
data.parquet
Full financial metrics (wide table, 100+ columns)
Key Fields (common)
Field… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_fin_indicators.indicvoices_r
IndicVoices-R: Multilingual, Multi-Speaker Speech Corpus for Indian TTS
Dataset Summary
IndicVoices-R (IV-R) is the largest multilingual Indian text-to-speech (TTS) dataset derived from an automatic speech recognition (ASR) dataset. It contains 1,704 hours of high-quality speech from 10,496 speakers across 22 Indian languages. This dataset is designed to enhance the development of robust Indian TTS models by providing diverse speaker demographics, natural… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/indicvoices_r.IndicSynth
IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset
A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research
🏆 Outstanding Paper Award, ACL 2025
🧠 Overview
IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/vdivyasharma/IndicSynth.NCERT-Parallel-Dataset-IndicIndicCorpV2
IndicCorp v2 Dataset
Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic Languages
This repository contains the pretraining data for the paper published at ACL 2023.
Example Usage
from datasets import load_dataset
# Load the Telugu subset of the dataset
dataset = load_dataset("ai4bharat/IndicCorpV2", "indiccorp_v2", data_dir="data/tel_Telu")
License
All the datasets created as… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicCorpV2.indic_glue
Dataset Card for "indic_glue"
Dataset Summary
IndicGLUE is a natural language understanding benchmark for Indian languages. It contains a wide
variety of tasks and covers 11 major Indian languages - as, bn, gu, hi, kn, ml, mr, or, pa, ta, te.
The Winograd Schema Challenge (Levesque et al., 2011) is a reading comprehension task
in which a system must read a sentence with a pronoun and select the referent of that pronoun from
a list of choices. The examples are manually… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/indic_glue.indic-dialect-asr
Indic Dialect ASR Dataset
A multilingual ASR dataset covering 30 Indic dialect/languages with 2.8M+ samples.
Usage
from datasets import load_dataset
# Load a specific language
ds = load_dataset("grushaaaaa/indic-dialect-asr", "assamese", split="train")
Features
audio: 16kHz WAV audio
sentence: Transcription text
language: Language name
source: Source dataset
IndicGenBenchFloresBitextMining
IndicGenBenchFloresBitextMining
An MTEB dataset
Massive Text Embedding Benchmark
Flores-IN dataset is an extension of Flores dataset released as a part of the IndicGenBench by Google
Task category
t2t
Domains
Web, News, Written
Reference
https://github.com/google-research-datasets/indic-gen-bench/
Source datasets:
google/IndicGenBench_flores_in
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:… See the full description on the dataset page: https://huggingface.co/datasets/mteb/IndicGenBenchFloresBitextMining.Indices-Daily-Price
Indices Daily Price
This dataset includes daily price data for various indices.
815,441 rows over 113 symbols, 8 columns, covering 1927-12-30 to 2026-08-03. Refreshed monthly.
Strategies Built on This Data
1,324 papers in the Papers With Backtest catalogue declare this dataset as an input. 1,226 of them have been coded and run over their own full history. The median replicated Sharpe ratio is +0.49, and 62% clear a t-statistic of 1.96 on their own sample, against… See the full description on the dataset page: https://huggingface.co/datasets/paperswithbacktest/Indices-Daily-Price.IndicSynth
IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset
A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research
🏆 Outstanding Paper Award, ACL 2025
🧠 Overview
IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/ksmashhero/IndicSynth.Finance-Conversational-Dataset-Indicindic-tts-966h
Indic-TTS-966h
Six-language Indian TTS corpus: ~966 hours of paired speech and text, 24 kHz mono WAV
clips with sentence-level transcripts in native scripts (natural English code-switching
preserved).
Subset
Clips
Hours
bengali
18,343
94.9
malayalam
30,548
192.5
marathi
34,327
213.4
punjabi
28,083
161.8
tamil
26,817
171.1
telugu
21,923
132.8
Columns: audio (24 kHz mono), file_name, transcript. One config per language:
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/psk/indic-tts-966h.indic-align
IndicAlign
A diverse collection of Instruction and Toxic alignment datasets for 14 Indic Languages. The collection comprises of:
IndicAlign - Instruct
Indic-ShareLlama
Dolly-T
OpenAssistant-T
WikiHow
IndoWordNet
Anudesh
Wiki-Conv
Wiki-Chat
IndicAlign - Toxic
HHRLHF-T
Toxic-Matrix
We use IndicTrans2 (Gala et al., 2023) for the translation of the datasets.
We recommend the readers to check out our paper on Arxiv for detailed information on the curation process of these… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/indic-align.indic-lma-corpus-phase1IndicParaphraseThis is the paraphrasing dataset released as part of IndicNLG Suite. Each
input is paired with up to 5 references. We create this dataset in eleven
languages including as, bn, gu, hi, kn, ml, mr, or, pa, ta, te. The total
size of the dataset is 5.57M.IndicWikiBioThis is the WikiBio dataset released as part of IndicNLG Suite. Each
example has four fields: id, infobox, serialized infobox and summary. We create this dataset in nine
languages including as, bn, hi, kn, ml, or, pa, ta, te. The total
size of the dataset is 57,426.IndicSentiment
IndicSentimentClassification
An MTEB dataset
Massive Text Embedding Benchmark
A new, multilingual, and n-way parallel dataset for sentiment analysis in 13 Indic languages.
Task category
t2c
Domains
Reviews, Written
Referencehttps://arxiv.org/abs/2212.05409
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["IndicSentimentClassification"])
evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/IndicSentiment.indic-diarbench
Indic DiarBench
A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio.
Paper: Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages (Interspeech 2026)
Dataset Summary
Indic DiarBench is a conversational speech benchmark designed to evaluate speaker-attributed ASR in realistic multi-speaker settings for… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indic-diarbench.Law-Conversational-Dataset-IndicCyber-Conversational-Dataset-IndicIndicTTS-Hindi
Hindi Indic TTS Dataset
This dataset is derived from the Indic TTS Database project, specifically using the Hindi monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development.
Dataset Details
Language: Hindi
Total Duration: ~10.33 hours (Male: 5.16 hours, Female: 5.18 hours)
Audio Format: WAV
Sampling Rate: 48000Hz… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS-Hindi.NCERT-Conversational-Dataset-IndicIndicTTS-Deepfake-Challenge-Data
IndicTTS Deepfake Detection Challenge
Participants will use the SherryT997/IndicTTS-Deepfake-Challenge-Data dataset, hosted on Hugging Face. This dataset consists of train and test splits and contains speech samples in 16 Indian languages, along with metadata for each audio clip.
🚀 Dataset to Use: SherryT997/IndicTTS-Deepfake-Challenge-Data
This is the official dataset for the challenge and must be used for training and evaluation.
📌 Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/SherryT997/IndicTTS-Deepfake-Challenge-Data.indicxnliIndicXNLI is a translated version of XNLI to 11 Indic Languages. As with XNLI, the goal is
to predict textual entailment (does sentence A imply/contradict/neither sentence
B) and is a classification task (given two sentences, predict one of three
labels).IndicTTS-EnglishMalayalam_CultureX_IndicCorp_SMCMalayalam Pretraining/Tokenization dataset.
Preprocessed and combined data from the following links,
* ai4bharat
* CulturaX
* Swathanthra Malayalam Computing
Commands used for preprocessing.
To remove all non Malayalam characters.
sed -i 's/[^ം-ൿ.,;:@$%+&?!() ]//g' test.txt
To merge all the text files in a particular Directory(Sub-Directory)
find SMC -type f -name '*.txt' -exec cat {} ; >> combined_SMC.txt
To remove all lines with characters less than 5.
grep -P… See the full description on the dataset page: https://huggingface.co/datasets/VishnuPJ/Malayalam_CultureX_IndicCorp_SMC.Environment-and-Natural-Resources-Indicators-For-African-Countries
Environment and Natural Resources Indicators For African Countries | Africa (World Health Organization)
Size category: 1K<n<10K - Formats: csv - Sector: climate_environment - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Environment-and-Natural-Resources-Indicators-For-African-Countries.Cyber-Parallel-Dataset-IndicIndicSynth
IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset
A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research
🏆 Outstanding Paper Award, ACL 2025
🧠 Overview
IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/mrunmai18/IndicSynth.
