CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ai4bharat /IndicVoicesgated IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages Updates [23 December 2025] We now have 11,200 hours of transcribed data! 🎉 Overview INDICVOICES is a dataset of natural and spontaneous speech containing a total of 23.7K hours of read (8%), extempore (76%) and conversational (15%) audio from 51K speakers covering 400+ Indian districts and 22 languages. Of these 23.7K hours, 11.2K hours have… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicVoices.audio1M<n<10M114 likes25k downloads3mo agoHugging Face02AlphaDojo /dojo_fin_indicators Languages: 简体中文 · English dojo_fin_indicators — Financial Metrics Overview Multi-period financial statement derivatives per symbol: income statement, balance sheet, cash flow, and industry-specific metrics (banks, insurers, brokers, etc.). Supports quarterly, cumulative, and other report_type values. Files File Description data.parquet Full financial metrics (wide table, 100+ columns) Key Fields (common) Field… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_fin_indicators.tabular100K<n<1M0 likes14k downloads10h agoHugging Face03ai4bharat /indicvoices_rgated IndicVoices-R: Multilingual, Multi-Speaker Speech Corpus for Indian TTS Dataset Summary IndicVoices-R (IV-R) is the largest multilingual Indian text-to-speech (TTS) dataset derived from an automatic speech recognition (ASR) dataset. It contains 1,704 hours of high-quality speech from 10,496 speakers across 22 Indian languages. This dataset is designed to enhance the development of robust Indian TTS models by providing diverse speaker demographics, natural… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/indicvoices_r.audiotext-to-speech100K<n<1M44 likes7.4k downloads2y agoHugging Face04vdivyasharma /IndicSynth IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research 🏆 Outstanding Paper Award, ACL 2025 🧠 Overview IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/vdivyasharma/IndicSynth.audioaudio-classification1M<n<10M14 likes5.9k downloads8mo agoHugging Face05oss-codes /NCERT-Parallel-Dataset-Indictexttranslation100K<n<1M2 likes4.6k downloads2y agoHugging Face06ai4bharat /IndicCorpV2 IndicCorp v2 Dataset Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic Languages This repository contains the pretraining data for the paper published at ACL 2023. Example Usage from datasets import load_dataset # Load the Telugu subset of the dataset dataset = load_dataset("ai4bharat/IndicCorpV2", "indiccorp_v2", data_dir="data/tel_Telu") License All the datasets created as… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicCorpV2.text100M<n<1B24 likes4.2k downloads2y agoHugging Face07ai4bharat /indic_glue Dataset Card for "indic_glue" Dataset Summary IndicGLUE is a natural language understanding benchmark for Indian languages. It contains a wide variety of tasks and covers 11 major Indian languages - as, bn, gu, hi, kn, ml, mr, or, pa, ta, te. The Winograd Schema Challenge (Levesque et al., 2011) is a reading comprehension task in which a system must read a sentence with a pronoun and select the referent of that pronoun from a list of choices. The examples are manually… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/indic_glue.tabulartext-classification100K<n<1M15 likes3.8k downloads3y agoHugging Face08grushaaaaa /indic-dialect-asr Indic Dialect ASR Dataset A multilingual ASR dataset covering 30 Indic dialect/languages with 2.8M+ samples. Usage from datasets import load_dataset # Load a specific language ds = load_dataset("grushaaaaa/indic-dialect-asr", "assamese", split="train") Features audio: 16kHz WAV audio sentence: Transcription text language: Language name source: Source dataset audioautomatic-speech-recognition1M<n<10M4 likes3.2k downloads7mo agoHugging Face09mteb /IndicGenBenchFloresBitextMining IndicGenBenchFloresBitextMining An MTEB dataset Massive Text Embedding Benchmark Flores-IN dataset is an extension of Flores dataset released as a part of the IndicGenBench by Google Task category t2t Domains Web, News, Written Reference https://github.com/google-research-datasets/indic-gen-bench/ Source datasets: google/IndicGenBench_flores_in How to evaluate on this task You can evaluate an embedding model on this dataset using the following code:… See the full description on the dataset page: https://huggingface.co/datasets/mteb/IndicGenBenchFloresBitextMining.texttranslation100K<n<1M1 likes3k downloads7mo agoHugging Face10paperswithbacktest /Indices-Daily-Pricegated Indices Daily Price This dataset includes daily price data for various indices. 815,441 rows over 113 symbols, 8 columns, covering 1927-12-30 to 2026-08-03. Refreshed monthly. Strategies Built on This Data 1,324 papers in the Papers With Backtest catalogue declare this dataset as an input. 1,226 of them have been coded and run over their own full history. The median replicated Sharpe ratio is +0.49, and 62% clear a t-statistic of 1.96 on their own sample, against… See the full description on the dataset page: https://huggingface.co/datasets/paperswithbacktest/Indices-Daily-Price.tabulartime-series-forecasting100K<n<1M2 likes2.9k downloads22d agoHugging Face11ksmashhero /IndicSynth IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research 🏆 Outstanding Paper Award, ACL 2025 🧠 Overview IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/ksmashhero/IndicSynth.audioaudio-classification1M<n<10M0 likes2.5k downloads14d agoHugging Face12oss-codes /Finance-Conversational-Dataset-Indictext100K<n<1M1 likes2.3k downloads1y agoHugging Face13psk /indic-tts-966h Indic-TTS-966h Six-language Indian TTS corpus: ~966 hours of paired speech and text, 24 kHz mono WAV clips with sentence-level transcripts in native scripts (natural English code-switching preserved). Subset Clips Hours bengali 18,343 94.9 malayalam 30,548 192.5 marathi 34,327 213.4 punjabi 28,083 161.8 tamil 26,817 171.1 telugu 21,923 132.8 Columns: audio (24 kHz mono), file_name, transcript. One config per language: from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/psk/indic-tts-966h.audio100K<n<1M6 likes2.2k downloads2mo agoHugging Face14ai4bharat /indic-align IndicAlign A diverse collection of Instruction and Toxic alignment datasets for 14 Indic Languages. The collection comprises of: IndicAlign - Instruct Indic-ShareLlama Dolly-T OpenAssistant-T WikiHow IndoWordNet Anudesh Wiki-Conv Wiki-Chat IndicAlign - Toxic HHRLHF-T Toxic-Matrix We use IndicTrans2 (Gala et al., 2023) for the translation of the datasets. We recommend the readers to check out our paper on Arxiv for detailed information on the curation process of these… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/indic-align.tabulartext-generation10M<n<100M21 likes2.1k downloads2y agoHugging Face15gurukondaveeti /indic-lma-corpus-phase1imagen<1K0 likes2k downloads8d agoHugging Face16ai4bharat /IndicParaphraseThis is the paraphrasing dataset released as part of IndicNLG Suite. Each input is paired with up to 5 references. We create this dataset in eleven languages including as, bn, gu, hi, kn, ml, mr, or, pa, ta, te. The total size of the dataset is 5.57M.text1M<n<10M6 likes1.9k downloads4y agoHugging Face17ai4bharat /IndicWikiBioThis is the WikiBio dataset released as part of IndicNLG Suite. Each example has four fields: id, infobox, serialized infobox and summary. We create this dataset in nine languages including as, bn, hi, kn, ml, or, pa, ta, te. The total size of the dataset is 57,426.text10K<n<100K2 likes1.8k downloads4y agoHugging Face18mteb /IndicSentiment IndicSentimentClassification An MTEB dataset Massive Text Embedding Benchmark A new, multilingual, and n-way parallel dataset for sentiment analysis in 13 Indic languages. Task category t2c Domains Reviews, Written Referencehttps://arxiv.org/abs/2212.05409 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["IndicSentimentClassification"]) evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/IndicSentiment.texttext-classification10K<n<100K0 likes1.5k downloads1y agoHugging Face19sarvamai /indic-diarbench Indic DiarBench A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio. Paper: Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages (Interspeech 2026) Dataset Summary Indic DiarBench is a conversational speech benchmark designed to evaluate speaker-attributed ASR in realistic multi-speaker settings for… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indic-diarbench.audioautomatic-speech-recognition1K<n<10K18 likes1.4k downloads1mo agoHugging Face20oss-codes /Law-Conversational-Dataset-Indictext100K<n<1M0 likes1.4k downloads1y agoHugging Face21oss-codes /Cyber-Conversational-Dataset-Indictext1K<n<10K0 likes1.3k downloads1y agoHugging Face22SPRINGLab /IndicTTS-Hindi Hindi Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Hindi monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development. Dataset Details Language: Hindi Total Duration: ~10.33 hours (Male: 5.16 hours, Female: 5.18 hours) Audio Format: WAV Sampling Rate: 48000Hz… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS-Hindi.audiotext-to-speech10K<n<100K36 likes1.3k downloads2y agoHugging Face23oss-codes /NCERT-Conversational-Dataset-Indictext100K<n<1M0 likes1.3k downloads1y agoHugging Face24SherryT997 /IndicTTS-Deepfake-Challenge-Data IndicTTS Deepfake Detection Challenge Participants will use the SherryT997/IndicTTS-Deepfake-Challenge-Data dataset, hosted on Hugging Face. This dataset consists of train and test splits and contains speech samples in 16 Indian languages, along with metadata for each audio clip. 🚀 Dataset to Use: SherryT997/IndicTTS-Deepfake-Challenge-Data This is the official dataset for the challenge and must be used for training and evaluation. 📌 Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/SherryT997/IndicTTS-Deepfake-Challenge-Data.audio10K<n<100K1 likes1.2k downloads2y agoHugging Face25Divyanshu /indicxnliIndicXNLI is a translated version of XNLI to 11 Indic Languages. As with XNLI, the goal is to predict textual entailment (does sentence A imply/contradict/neither sentence B) and is a classification task (given two sentences, predict one of three labels).texttext-classification1M<n<10M7 likes1.2k downloads4y agoHugging Face26SPRINGLab /IndicTTS-Englishaudio100K<n<1M2 likes1.2k downloads2y agoHugging Face27VishnuPJ /Malayalam_CultureX_IndicCorp_SMCMalayalam Pretraining/Tokenization dataset. Preprocessed and combined data from the following links, * ai4bharat * CulturaX * Swathanthra Malayalam Computing Commands used for preprocessing. To remove all non Malayalam characters. sed -i 's/[^ം-ൿ.,;:@$%+&?!() ]//g' test.txt To merge all the text files in a particular Directory(Sub-Directory) find SMC -type f -name '*.txt' -exec cat {} ; >> combined_SMC.txt To remove all lines with characters less than 5. grep -P… See the full description on the dataset page: https://huggingface.co/datasets/VishnuPJ/Malayalam_CultureX_IndicCorp_SMC.text10M<n<100M4 likes1.1k downloads2y agoHugging Face28electricsheepafrica /Environment-and-Natural-Resources-Indicators-For-African-Countries Environment and Natural Resources Indicators For African Countries | Africa (World Health Organization) Size category: 1K<n<10K - Formats: csv - Sector: climate_environment - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Environment-and-Natural-Resources-Indicators-For-African-Countries.tabulartabular-classification1K<n<10K0 likes1.1k downloads1mo agoHugging Face29oss-codes /Cyber-Parallel-Dataset-Indictext1K<n<10K0 likes1k downloads1y agoHugging Face30mrunmai18 /IndicSynth IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research 🏆 Outstanding Paper Award, ACL 2025 🧠 Overview IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/mrunmai18/IndicSynth.audioaudio-classification1M<n<10M0 likes985 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.