datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Finance-Conversational-Dataset-IndicLaw-Conversational-Dataset-IndicCyber-Conversational-Dataset-IndicNCERT-Conversational-Dataset-IndicIndicSentiment
IndicSentimentClassification
An MTEB dataset
Massive Text Embedding Benchmark
A new, multilingual, and n-way parallel dataset for sentiment analysis in 13 Indic languages.
Task category
t2c
Domains
Reviews, Written
Referencehttps://arxiv.org/abs/2212.05409
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["IndicSentimentClassification"])
evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/IndicSentiment.Cyber-Parallel-Dataset-IndicIndicVoice-latent-NEWFinance-Parallel-Dataset-IndicComputer-Science-Conversational-Dataset-IndicLaw-Parallel-Dataset-IndicCA-Conversational-Dataset-IndicIndicNLP-MultilingualMedical-Parallel-Dataset-Indicindicphi
IndicPHI
Synthetic Indian clinical documents with PHI/PII span labels for 23 languages (22 scheduled Indian languages + English). All identifiers are synthetic surrogates, not real patients.
Two Hub configs at the repo root:
Config
Files
Rows
What it is
default
train.jsonl, eval.jsonl
22,554 / 3,982
Full documents: text, character spans, metadata
gliner
gliner_train.json, gliner_eval.json
22,889 / 4,054
GLiNER windows: tokenized_text + token ner
Also on the… See the full description on the dataset page: https://huggingface.co/datasets/Sidharth1743/indicphi.indic-queries-2026
MAST Indic Queries 2026
This dataset contains the Indic query set for MAST @ FIRE 2026, the Multilingual Agentic Search Track. MAST evaluates whether multilingual agentic search systems can answer complex questions posed in different languages.
MAST builds on BrowseComp-Plus (ACL 2026), a reproducible and verifiable extension of BrowseComp with challenging English queries, a verified English corpus of roughly 100K web-sourced documents, and human judgments. In the 2026 MAST… See the full description on the dataset page: https://huggingface.co/datasets/mast-benchmark/indic-queries-2026.Medical-Conversational-Dataset-IndicCA-Parallel-Dataset-Indicindic-squad
IndicSQuAD Dataset
Dataset Description
IndicSQuAD is a comprehensive multilingual extractive Question Answering (QA) dataset covering nine major Indic languages: Hindi, Bengali, Tamil, Telugu, Marathi, Gujarati, Urdu, Kannada, Oriya, and Malayalam. It's systematically derived from the popular English SQuAD (Stanford Question Answering Dataset).
The rapid progress in QA systems has predominantly benefited high-resource languages, leaving Indic languages significantly… See the full description on the dataset page: https://huggingface.co/datasets/l3cube-pune/indic-squad.CAT-Conversational-Dataset-IndicConstrained_Indic_CodemixingIndicTalk
IndicTalk:Code-Mixed Conversational Persona-Based Dataset
This dataset contains multi-turn, persona-driven, code-mixed conversations
generated from real news articles, across 9 Indian
languages, in two script variants:
Native — conversations written in the language's native script,
code-mixed with Romanized English words.
Romanized — conversations fully Romanized (Latin script), code-mixed
with English.
Each language has its own config, loadable independently, e.g.:
from… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/IndicTalk.indic-instruct-data-v0.1-filteredThis is filtered version of indic-instruct-data-v0.1.UPDATE: 4 March 2024 - This dataset has been further filtered to create indic-instruct-data-v0.2-filtered.
Filtering Approach
Drop exampels containing ["search the web", "www.", ".py", ".com", "spanish", "french", "japanese", "given two strings, check whether one string is a rotation of another", "openai", "xml", "arrange the words", "__", "noinput" "idiom", "alphabetic", "alliteration", "translat", "paraphrase", "code", "def "… See the full description on the dataset page: https://huggingface.co/datasets/BhabhaAI/indic-instruct-data-v0.1-filtered.indic-instruct-data-v0.2-filteredThis is v0.2 of indic-instruct-data-v0.1-filteredNote: lmsys dataset contain NAME_1, NAME_2 etc. You may replace them with actual names before fine-tuning.
Jee-Parallel-Dataset-Indicindice-legal-colombia
Indice legal de Colombia - Aliado Libre
Indice de legislacion y jurisprudencia colombiana: 718.388 fragmentos de
119.708 documentos oficiales, construido por
Aliado Libre - RAG legal
gratuito y de codigo abierto.
Son dos piezas, y la busqueda necesita las dos:
Archivo
Que es
Tamano
chroma.sqlite3 + carpetas UUID
Indice vectorial ChromaDB
~19 GB
fts_index.db
Indice lexico SQLite FTS5 (BM25 en disco)
~1,6 GB
Que embeddings tiene, y por que hay dos… See the full description on the dataset page: https://huggingface.co/datasets/Gullax/indice-legal-colombia.PI-Indic-Align
PI-Indic-Align: Persona-Instruction Alignment for Indian Languages
🦚 PI-Indic-Align: Persona-Instruction Alignment for Indian Languages
Teaching AI to speak the languages of India, one persona at a time!
Dataset Description
PI-Indic-Align is a large-scale benchmark dataset for evaluating persona-instruction alignment across 12 major Indian languages. The dataset contains 600,000 culturally grounded persona-instruction pairs (50,000 per language) designed… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/PI-Indic-Align.IndicVault
Indic Vault — everyday Indian language QA pairs, tuned for chatbots & voice agents.
🧾 Overview
Indic Vault is a high-quality, instruction-tuned dataset featuring question-answer pairs crafted in the contemporary, everyday language spoken across India in 2025. Unlike traditional datasets that lean heavily on formal or outdated linguistic styles, Indic Vault captures the authentic, colloquial expressions used in daily conversations, making it ideal for building AI… See the full description on the dataset page: https://huggingface.co/datasets/maya-research/IndicVault.IndicBankBench
IndicBankBench — A Benchmark for Evaluating the Safety and Reliability of Language Models in Indian Retail Banking
IndicBankBench evaluates whether language models behave safely and reliably in Indian
retail-banking interactions. Its 799 synthetic, multi-turn cases test grounding in customer
context, safe action-taking with mocked banking tools, appropriate clarification and refusal,
and complete resolution of customer requests. All data is synthetic and contains no real… See the full description on the dataset page: https://huggingface.co/datasets/NPCI/IndicBankBench.IndicConBench
IndicConBench: A Multilingual Constitutional Reasoning Benchmark for India
IndicConBench is a publication-grade, research-quality multilingual constitutional reasoning benchmark for India. It is specifically designed to evaluate the factual accuracy, reasoning capacity, structural retrieval capability, and legal intelligence of Large Language Models (LLMs) on the Constitution of India in both English and Hindi.
Developed by Vikhram S and published under Vikhram Labs, this… See the full description on the dataset page: https://huggingface.co/datasets/Vikhram-S/IndicConBench.Indic-KCC-Agri-Advisory-Benchmark
Indic-KCC-Agri-Advisory-Benchmark
⚠️ Benchmark only — not agronomic advice. This dataset and its reference
answers exist to score language models, not to be used as real farming
guidance. KCC references are noisy call-centre transcripts (see Status and
caveats); do not act on any answer, reference or candidate, as agricultural
advice.
Open-ended agricultural-advisory question answering in 11 Indian languages,
built from real farmer questions and the advisory answers… See the full description on the dataset page: https://huggingface.co/datasets/sthanika-ai/Indic-KCC-Agri-Advisory-Benchmark.
