CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ai4bharat /IndicVoicesgated IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages Updates [23 December 2025] We now have 11,200 hours of transcribed data! 🎉 Overview INDICVOICES is a dataset of natural and spontaneous speech containing a total of 23.7K hours of read (8%), extempore (76%) and conversational (15%) audio from 51K speakers covering 400+ Indian districts and 22 languages. Of these 23.7K hours, 11.2K hours have… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicVoices.audio1M<n<10M115 likes25k downloads3mo agoHugging Face02vaquill /open-india-lawgated Open India Law Open, structured Indian primary law - plus the scrapers that build it. Every judgment of the Supreme Court of India and all 25 High Courts, the decisions of 15 tribunals and regulators, and Central, State and Union Territory legislation down to the individual section. Normalized to one schema, exclusively from official government sources. Volume Period Court judgments 12,848,644 1950 to 2025 Tribunal and regulator matters 813,168 1985 to 2026… See the full description on the dataset page: https://huggingface.co/datasets/vaquill/open-india-law.tabulartext-retrieval10M<n<100M26 likes14k downloads1mo agoHugging Face03AlphaDojo /dojo_fin_indicators Languages: 简体中文 · English dojo_fin_indicators — Financial Metrics Overview Multi-period financial statement derivatives per symbol: income statement, balance sheet, cash flow, and industry-specific metrics (banks, insurers, brokers, etc.). Supports quarterly, cumulative, and other report_type values. Files File Description data.parquet Full financial metrics (wide table, 100+ columns) Key Fields (common) Field… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_fin_indicators.tabular100K<n<1M0 likes14k downloads2h agoHugging Face04mratanusarkar /Indian-Laws Dataset Card for Indian Laws This is a comprehensive collection of primary legal documents pertinent to the Indian legal system. It is designed to serve as a foundational resource for supervised fine-tuning (SFT) to make language models, particularly those focused on legal applications tailored for Indian law. text10K<n<100K8 likes12k downloads3y agoHugging Face05thetrademarkk /india-index-options-1m India Index & Options - 1-minute OHLC 1-minute OHLCV(+OI) bars for NSE/BSE index spot and option chains: NIFTY, BANKNIFTY, SENSEX (~2021-2026). Powers the open-source TradeMarkk backtester (https://thetrademarkk.com). Educational use only. Provided as-is, no warranty. Verify against official exchange data before relying on it. Structure index/{SYMBOL}.parquet - 1-min spot OHLC per index. options/{SYMBOL}/{EXPIRY}.parquet - 1-min OHLC per option contract (with… See the full description on the dataset page: https://huggingface.co/datasets/thetrademarkk/india-index-options-1m.tabular100M<n<1B4 likes8.8k downloads3mo agoHugging Face06wayu-ai /thai-commoncrawl-index Thai Common Crawl Index (2019–2026) An index of every page Common Crawl detected as Thai across 70 monthly crawls, from January 2019 (CC-MAIN-2019-04) to August 2026 (CC-MAIN-2026-30). 932,874,727 page captures · 450,971,497 unique URLs · 6,997,185 hosts · 6,674,969 domains Each row records where the page lives inside Common Crawl's WARC archives — file name, byte offset, and record length — so you can fetch exactly the pages you want with HTTP range requests, without scanning… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-commoncrawl-index.tabular100M<n<1B0 likes8.1k downloads1mo agoHugging Face07ai4bharat /indicvoices_rgated IndicVoices-R: Multilingual, Multi-Speaker Speech Corpus for Indian TTS Dataset Summary IndicVoices-R (IV-R) is the largest multilingual Indian text-to-speech (TTS) dataset derived from an automatic speech recognition (ASR) dataset. It contains 1,704 hours of high-quality speech from 10,496 speakers across 22 Indian languages. This dataset is designed to enhance the development of robust Indian TTS models by providing diverse speaker demographics, natural… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/indicvoices_r.audiotext-to-speech100K<n<1M44 likes7.8k downloads2y agoHugging Face08commoncrawl /host-index-testing-v2 Common Crawl Host Index v2 GitHub: https://github.com/commoncrawl/cc-host-index Each crawl, we generate a Host Index, which aggregates information about each web hosted visited during the crawl. The information is aggregated from the Common Crawl columnar index, web graph, and raw crawler logs. Quickstart The dataset is Hive-partitioned on crawl (data/crawl=CC-MAIN-2025-18/*.parquet). Open the whole dataset once, then filter with WHERE crawl = '...': because… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/host-index-testing-v2.tabulartext-generation1B<n<10B0 likes7.2k downloads11d agoHugging Face09jjldo21 /IndustrialDetectionStaticCamerasThe IndustrialDetectionStaticCameras dataset has been collected in order to validate the methodology presented in the paper entitled A few-shot learning methodology for improving safety in industrial scenarios through universal self-supervised visual features and dense optical flow. This dataset is divided into five main folders named videoY, where Y=1,2,3,4,5. Each videoY folder contains the following: The video of the scene in .mp4 format: videoY.mp4 A folder with the images of each frame… See the full description on the dataset page: https://huggingface.co/datasets/jjldo21/IndustrialDetectionStaticCameras.imageobject-detection1K<n<10K1 likes6.1k downloads2y agoHugging Face10gagandeepreehal /minuszero-indian-autonomous-driving-dataset-v2gated INDUS-AD: Indian Dataset of Unstructured Urban Scenes for Autonomous Driving Overview INDUS-AD is the largest publicly released Indian autonomous-driving dataset for end-to-end autonomous-driving research. Its name expands to Indian Dataset of Unstructured Urban Scenes for Autonomous Driving. This gated dataset is the decoded companion to the Minus Zero Indian Urban Autonomous Driving Dataset. It provides directly usable camera MP4s, normalized sensor tables… See the full description on the dataset page: https://huggingface.co/datasets/gagandeepreehal/minuszero-indian-autonomous-driving-dataset-v2.imagerobotics10M<n<100M1 likes5.4k downloads7d agoHugging Face11vdivyasharma /IndicSynth IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research 🏆 Outstanding Paper Award, ACL 2025 🧠 Overview IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/vdivyasharma/IndicSynth.audioaudio-classification1M<n<10M14 likes5.2k downloads9mo agoHugging Face12Vikaschou /Indian-Laws Dataset Card for Indian Laws This is a comprehensive collection of primary legal documents pertinent to the Indian legal system. It is designed to serve as a foundational resource for supervised fine-tuning (SFT) to make language models, particularly those focused on legal applications tailored for Indian law. text10K<n<100K1 likes4.7k downloads6mo agoHugging Face13oss-codes /NCERT-Parallel-Dataset-Indictexttranslation100K<n<1M2 likes4.4k downloads2y agoHugging Face14ai4bharat /indic_glue Dataset Card for "indic_glue" Dataset Summary IndicGLUE is a natural language understanding benchmark for Indian languages. It contains a wide variety of tasks and covers 11 major Indian languages - as, bn, gu, hi, kn, ml, mr, or, pa, ta, te. The Winograd Schema Challenge (Levesque et al., 2011) is a reading comprehension task in which a system must read a sentence with a pronoun and select the referent of that pronoun from a list of choices. The examples are manually… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/indic_glue.tabulartext-classification100K<n<1M15 likes4.2k downloads3y agoHugging Face15physicl /indoor-safety-hazard-detection-and-work-zone-monitoring Indoor Safety Hazard Detection & Work-Zone Monitoring Generated by datapack-import.ts This dataset mirrors public data-pack render outputs from Physicl. Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by… See the full description on the dataset page: https://huggingface.co/datasets/physicl/indoor-safety-hazard-detection-and-work-zone-monitoring.imagen<1K0 likes4.1k downloads3mo agoHugging Face16BAAI /IndustryCorpus2_tourism_geography IndustryCorpus2: Travel & Geography This repository contains the IndustryCorpus2: Travel & Geography domain subset of BAAI/IndustryCorpus2. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryCorpus2: @misc{shi2024industrycorpus2, title = {IndustryCorpus2}, author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao}, year… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_tourism_geography.tabular10M<n<100M4 likes3.8k downloads1mo agoHugging Face17vidore /vidore_v3_industrialViDoRe V3 : Industrial reports This dataset, Industrial reports, is a corpus of technical documents on military aircrafts (fueling, mechanics...), intended for complex-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark. About ViDoRe v3 ViDoRe V3 is our latest benchmark for RAG evaluation on visually-rich documents from real-world applications. It features 10 datasets with, in total, 26,000 pages and 3099 queries, translated into 6… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_industrial.documentvisual-document-retrieval10K<n<100K7 likes3.7k downloads8mo agoHugging Face18grushaaaaa /indic-dialect-asr Indic Dialect ASR Dataset A multilingual ASR dataset covering 30 Indic dialect/languages with 2.8M+ samples. Usage from datasets import load_dataset # Load a specific language ds = load_dataset("grushaaaaa/indic-dialect-asr", "assamese", split="train") Features audio: 16kHz WAV audio sentence: Transcription text language: Language name source: Source dataset audioautomatic-speech-recognition1M<n<10M4 likes3.6k downloads8mo agoHugging Face19Becky7777777 /polymarket-search-indextabular1M<n<10M0 likes3.6k downloads15h agoHugging Face20kumarmanoj382 /Inddatatext1B<n<10B0 likes3.4k downloads2mo agoHugging Face21ksmashhero /IndicSynth IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research 🏆 Outstanding Paper Award, ACL 2025 🧠 Overview IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/ksmashhero/IndicSynth.audioaudio-classification1M<n<10M0 likes3.4k downloads17d agoHugging Face22paperswithbacktest /Indices-Daily-Pricegated Indices Daily Price This dataset includes daily price data for various indices. 815,441 rows over 113 symbols, 8 columns, covering 1927-12-30 to 2026-08-03. Refreshed monthly. Strategies Built on This Data 1,324 papers in the Papers With Backtest catalogue declare this dataset as an input. 1,226 of them have been coded and run over their own full history. The median replicated Sharpe ratio is +0.49, and 62% clear a t-statistic of 1.96 on their own sample, against… See the full description on the dataset page: https://huggingface.co/datasets/paperswithbacktest/Indices-Daily-Price.tabulartime-series-forecasting100K<n<1M2 likes3.3k downloads24d agoHugging Face23open-index /umi-robots umi-robots Part of umi, an open web crawl published as Parquet. Before you use any of this, read the exclusion list at open-index/umi-meta and filter the rows it names. Published files are never rewritten, so the exclusion list is how a takedown reaches you, and applying it is a condition of using the data rather than a suggestion. One row per robots.txt fetch: the host, when we asked, what the origin answered, the raw text if it served one, and the summary our parser read out… See the full description on the dataset page: https://huggingface.co/datasets/open-index/umi-robots.tabular10M<n<100M0 likes3.2k downloads19d agoHugging Face24mteb /IndicGenBenchFloresBitextMining IndicGenBenchFloresBitextMining An MTEB dataset Massive Text Embedding Benchmark Flores-IN dataset is an extension of Flores dataset released as a part of the IndicGenBench by Google Task category t2t Domains Web, News, Written Reference https://github.com/google-research-datasets/indic-gen-bench/ Source datasets: google/IndicGenBench_flores_in How to evaluate on this task You can evaluate an embedding model on this dataset using the following code:… See the full description on the dataset page: https://huggingface.co/datasets/mteb/IndicGenBenchFloresBitextMining.texttranslation100K<n<1M1 likes3k downloads7mo agoHugging Face25PlumCascade /Equal_Industry_day_10_pqttabular10M<n<100M0 likes2.6k downloads5mo agoHugging Face26manjot007 /Indian-Laws Dataset Card for Indian Laws This is a comprehensive collection of primary legal documents pertinent to the Indian legal system. It is designed to serve as a foundational resource for supervised fine-tuning (SFT) to make language models, particularly those focused on legal applications tailored for Indian law. text10K<n<100K0 likes2.5k downloads5mo agoHugging Face27psk /indic-tts-966h Indic-TTS-966h Six-language Indian TTS corpus: ~966 hours of paired speech and text, 24 kHz mono WAV clips with sentence-level transcripts in native scripts (natural English code-switching preserved). Subset Clips Hours bengali 18,343 94.9 malayalam 30,548 192.5 marathi 34,327 213.4 punjabi 28,083 161.8 tamil 26,817 171.1 telugu 21,923 132.8 Columns: audio (24 kHz mono), file_name, transcript. One config per language: from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/psk/indic-tts-966h.audio100K<n<1M6 likes2.5k downloads2mo agoHugging Face28BAAI /IndustryCorpus2_medicine_health_psychology_traditional_chinese_medicine IndustryCorpus2: Health & Medicine This repository contains the IndustryCorpus2: Health & Medicine domain subset of BAAI/IndustryCorpus2. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryCorpus2: @misc{shi2024industrycorpus2, title = {IndustryCorpus2}, author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao}, year… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_medicine_health_psychology_traditional_chinese_medicine.tabular10M<n<100M11 likes2.3k downloads1mo agoHugging Face29tfqdeadlo /Inddatatext1B<n<10B0 likes2.1k downloads5mo agoHugging Face30ai4bharat /indic-align IndicAlign A diverse collection of Instruction and Toxic alignment datasets for 14 Indic Languages. The collection comprises of: IndicAlign - Instruct Indic-ShareLlama Dolly-T OpenAssistant-T WikiHow IndoWordNet Anudesh Wiki-Conv Wiki-Chat IndicAlign - Toxic HHRLHF-T Toxic-Matrix We use IndicTrans2 (Gala et al., 2023) for the translation of the datasets. We recommend the readers to check out our paper on Arxiv for detailed information on the curation process of these… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/indic-align.tabulartext-generation10M<n<100M21 likes2.1k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.