CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01oss-codes /Finance-Conversational-Dataset-Indictext100K<n<1M1 likes2.3k downloads1y agoHugging Face02oss-codes /Law-Conversational-Dataset-Indictext100K<n<1M0 likes1.4k downloads1y agoHugging Face03oss-codes /Cyber-Conversational-Dataset-Indictext1K<n<10K0 likes1.3k downloads1y agoHugging Face04oss-codes /NCERT-Conversational-Dataset-Indictext100K<n<1M0 likes1.3k downloads1y agoHugging Face05mteb /IndicSentiment IndicSentimentClassification An MTEB dataset Massive Text Embedding Benchmark A new, multilingual, and n-way parallel dataset for sentiment analysis in 13 Indic languages. Task category t2c Domains Reviews, Written Referencehttps://arxiv.org/abs/2212.05409 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["IndicSentimentClassification"]) evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/IndicSentiment.texttext-classification10K<n<100K0 likes1.2k downloads1y agoHugging Face06oss-codes /Cyber-Parallel-Dataset-Indictext1K<n<10K0 likes1.1k downloads1y agoHugging Face07zuhri025 /IndicVoice-latent-NEWtext100K<n<1M0 likes905 downloads5mo agoHugging Face08oss-codes /Finance-Parallel-Dataset-Indictext100K<n<1M0 likes825 downloads1y agoHugging Face09oss-codes /Computer-Science-Conversational-Dataset-Indictext10K<n<100K0 likes728 downloads1y agoHugging Face10oss-codes /Law-Parallel-Dataset-Indictext100K<n<1M0 likes636 downloads1y agoHugging Face11oss-codes /CA-Conversational-Dataset-Indictext100K<n<1M0 likes627 downloads1y agoHugging Face12Sakshamrzt /IndicNLP-Multilingualtexttext-classification10K<n<100K1 likes331 downloads2y agoHugging Face13oss-codes /Medical-Parallel-Dataset-Indictext10K<n<100K0 likes329 downloads1y agoHugging Face14Sidharth1743 /indicphi IndicPHI Synthetic Indian clinical documents with PHI/PII span labels for 23 languages (22 scheduled Indian languages + English). All identifiers are synthetic surrogates, not real patients. Two Hub configs at the repo root: Config Files Rows What it is default train.jsonl, eval.jsonl 22,554 / 3,982 Full documents: text, character spans, metadata gliner gliner_train.json, gliner_eval.json 22,889 / 4,054 GLiNER windows: tokenized_text + token ner Also on the… See the full description on the dataset page: https://huggingface.co/datasets/Sidharth1743/indicphi.texttoken-classification10K<n<100K1 likes302 downloads10d agoHugging Face15mast-benchmark /indic-queries-2026 MAST Indic Queries 2026 This dataset contains the Indic query set for MAST @ FIRE 2026, the Multilingual Agentic Search Track. MAST evaluates whether multilingual agentic search systems can answer complex questions posed in different languages. MAST builds on BrowseComp-Plus (ACL 2026), a reproducible and verifiable extension of BrowseComp with challenging English queries, a verified English corpus of roughly 100K web-sourced documents, and human judgments. In the 2026 MAST… See the full description on the dataset page: https://huggingface.co/datasets/mast-benchmark/indic-queries-2026.textquestion-answeringn<1K1 likes278 downloads2mo agoHugging Face16oss-codes /Medical-Conversational-Dataset-Indictext10K<n<100K0 likes262 downloads1y agoHugging Face17oss-codes /CA-Parallel-Dataset-Indictext100K<n<1M0 likes243 downloads1y agoHugging Face18l3cube-pune /indic-squad IndicSQuAD Dataset Dataset Description IndicSQuAD is a comprehensive multilingual extractive Question Answering (QA) dataset covering nine major Indic languages: Hindi, Bengali, Tamil, Telugu, Marathi, Gujarati, Urdu, Kannada, Oriya, and Malayalam. It's systematically derived from the popular English SQuAD (Stanford Question Answering Dataset). The rapid progress in QA systems has predominantly benefited high-resource languages, leaving Indic languages significantly… See the full description on the dataset page: https://huggingface.co/datasets/l3cube-pune/indic-squad.textquestion-answering1M<n<10M0 likes209 downloads1y agoHugging Face19oss-codes /CAT-Conversational-Dataset-Indictext10K<n<100K0 likes205 downloads1y agoHugging Face20Anvesh-Lankala /Constrained_Indic_Codemixingtext1K<n<10K0 likes165 downloads1mo agoHugging Face21LingoIITGN /IndicTalk IndicTalk:Code-Mixed Conversational Persona-Based Dataset This dataset contains multi-turn, persona-driven, code-mixed conversations generated from real news articles, across 9 Indian languages, in two script variants: Native — conversations written in the language's native script, code-mixed with Romanized English words. Romanized — conversations fully Romanized (Latin script), code-mixed with English. Each language has its own config, loadable independently, e.g.: from… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/IndicTalk.texttext-generation1M<n<10M2 likes159 downloads2mo agoHugging Face22BhabhaAI /indic-instruct-data-v0.1-filteredThis is filtered version of indic-instruct-data-v0.1.UPDATE: 4 March 2024 - This dataset has been further filtered to create indic-instruct-data-v0.2-filtered. Filtering Approach Drop exampels containing ["search the web", "www.", ".py", ".com", "spanish", "french", "japanese", "given two strings, check whether one string is a rotation of another", "openai", "xml", "arrange the words", "__", "noinput" "idiom", "alphabetic", "alliteration", "translat", "paraphrase", "code", "def "… See the full description on the dataset page: https://huggingface.co/datasets/BhabhaAI/indic-instruct-data-v0.1-filtered.text100K<n<1M0 likes158 downloads3y agoHugging Face23BhabhaAI /indic-instruct-data-v0.2-filteredThis is v0.2 of indic-instruct-data-v0.1-filteredNote: lmsys dataset contain NAME_1, NAME_2 etc. You may replace them with actual names before fine-tuning. text100K<n<1M0 likes157 downloads3y agoHugging Face24oss-codes /Jee-Parallel-Dataset-Indictext10K<n<100K0 likes128 downloads1y agoHugging Face25Gullax /indice-legal-colombia Indice legal de Colombia - Aliado Libre Indice de legislacion y jurisprudencia colombiana: 718.388 fragmentos de 119.708 documentos oficiales, construido por Aliado Libre - RAG legal gratuito y de codigo abierto. Son dos piezas, y la busqueda necesita las dos: Archivo Que es Tamano chroma.sqlite3 + carpetas UUID Indice vectorial ChromaDB ~19 GB fts_index.db Indice lexico SQLite FTS5 (BM25 en disco) ~1,6 GB Que embeddings tiene, y por que hay dos… See the full description on the dataset page: https://huggingface.co/datasets/Gullax/indice-legal-colombia.text100K<n<1M0 likes122 downloads17d agoHugging Face26LingoIITGN /PI-Indic-Align PI-Indic-Align: Persona-Instruction Alignment for Indian Languages 🦚 PI-Indic-Align: Persona-Instruction Alignment for Indian Languages Teaching AI to speak the languages of India, one persona at a time! Dataset Description PI-Indic-Align is a large-scale benchmark dataset for evaluating persona-instruction alignment across 12 major Indian languages. The dataset contains 600,000 culturally grounded persona-instruction pairs (50,000 per language) designed… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/PI-Indic-Align.texttext-retrieval1M<n<10M0 likes114 downloads8mo agoHugging Face27maya-research /IndicVault Indic Vault — everyday Indian language QA pairs, tuned for chatbots & voice agents. 🧾 Overview Indic Vault is a high-quality, instruction-tuned dataset featuring question-answer pairs crafted in the contemporary, everyday language spoken across India in 2025. Unlike traditional datasets that lean heavily on formal or outdated linguistic styles, Indic Vault captures the authentic, colloquial expressions used in daily conversations, making it ideal for building AI… See the full description on the dataset page: https://huggingface.co/datasets/maya-research/IndicVault.textquestion-answering100K<n<1M70 likes110 downloads1y agoHugging Face28NPCI /IndicBankBench IndicBankBench — A Benchmark for Evaluating the Safety and Reliability of Language Models in Indian Retail Banking IndicBankBench evaluates whether language models behave safely and reliably in Indian retail-banking interactions. Its 799 synthetic, multi-turn cases test grounding in customer context, safe action-taking with mocked banking tools, appropriate clarification and refusal, and complete resolution of customer requests. All data is synthetic and contains no real… See the full description on the dataset page: https://huggingface.co/datasets/NPCI/IndicBankBench.texttext-generationn<1K0 likes102 downloads2d agoHugging Face29Vikhram-S /IndicConBench IndicConBench: A Multilingual Constitutional Reasoning Benchmark for India IndicConBench is a publication-grade, research-quality multilingual constitutional reasoning benchmark for India. It is specifically designed to evaluate the factual accuracy, reasoning capacity, structural retrieval capability, and legal intelligence of Large Language Models (LLMs) on the Constitution of India in both English and Hindi. Developed by Vikhram S and published under Vikhram Labs, this… See the full description on the dataset page: https://huggingface.co/datasets/Vikhram-S/IndicConBench.textquestion-answering1K<n<10K0 likes101 downloads20d agoHugging Face30sthanika-ai /Indic-KCC-Agri-Advisory-Benchmarkgated Indic-KCC-Agri-Advisory-Benchmark ⚠️ Benchmark only — not agronomic advice. This dataset and its reference answers exist to score language models, not to be used as real farming guidance. KCC references are noisy call-centre transcripts (see Status and caveats); do not act on any answer, reference or candidate, as agricultural advice. Open-ended agricultural-advisory question answering in 11 Indian languages, built from real farmer questions and the advisory answers… See the full description on the dataset page: https://huggingface.co/datasets/sthanika-ai/Indic-KCC-Agri-Advisory-Benchmark.texttext-generation1K<n<10K4 likes99 downloads2d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.