CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-index /arctic Arctic Shift Reddit Archive Every Reddit comment and submission since 2005, organized as monthly Parquet shards What is it? The full Reddit archive from Arctic Shift, converted to Parquet and hosted here for easy access. Covers every public subreddit from 2005-12 through 2026-02. Right now the archive has 15.7B items (12.9B comments, 2.8B submissions) in 1.3 TB of compressed Parquet. Comments and submissions are stored as separate datasets, split into monthly… See the full description on the dataset page: https://huggingface.co/datasets/open-index/arctic.text-generation1B<n<10B29 likes100k downloads2mo agoHugging Face02deepghs /character_index Anime Character Index This dataset if for collecting all the hot characters from the internet, and extract their features and core tags. It will be useful for automatically testing the character generating ability of the anime-style base models. 7371 characters in total. Copyrights Copyright Count kantai_collection 393 pokemon 380 fate_(series) 350 hololive 277 blue_archive234 arknights 200 idolmaster 192 touhou 186 fire_emblem 168 umamusume… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/character_index.24 likes66k downloads1y agoHugging Face03open-index /hacker-news Hacker News - Complete Archive Every Hacker News item since 2006, live-updated every 5 minutes What is it? This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers… See the full description on the dataset page: https://huggingface.co/datasets/open-index/hacker-news.texttext-generation10M<n<100M343 likes57k downloads1mo agoHugging Face04FraunhoferIPK /IndEgo IndEgo: A Dataset of Industrial Scenarios and Collaborative Work for Egocentric Assistants Vivek Chavan¹²*, Yasmina Imgrund²†, Tung Dao²†, Sanwantri Bai³†, Bosong Wang⁴†, Ze Lu⁵†, Oliver Heimann¹, Jörg Krüger¹² ¹Fraunhofer IPK, Berlin &nbsp;&nbsp; ²Technical University of Berlin &nbsp;&nbsp; ³University of Tübingen ⁴RWTH Aachen University &nbsp;&nbsp; ⁵Leibniz University Hannover *Project Lead &nbsp;&nbsp;&nbsp; †Work done during student theses/projects at Fraunhofer IPK… See the full description on the dataset page: https://huggingface.co/datasets/FraunhoferIPK/IndEgo.visual-question-answering10K<n<100K9 likes38k downloads3mo agoHugging Face05fineweb-retrieval /fineweb-edu-indexThis dataset contains the embeddings for the full fineweb-edu, embedded with the Cohere Embed V3 model. You can search on this dataset with just 500MB of memory using DiskVectorIndex. Installation & Usage Get your free Cohere API key from cohere.com. You must set this API key as an environment variable: export COHERE_API_KEY=your_api_key Install the package: pip install DiskVectorIndex You can then search via: from DiskVectorIndex import DiskVectorIndex index =… See the full description on the dataset page: https://huggingface.co/datasets/fineweb-retrieval/fineweb-edu-index.0 likes29k downloads1y agoHugging Face06ai4bharat /IndicVoicesgated IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages Updates [23 December 2025] We now have 11,200 hours of transcribed data! 🎉 Overview INDICVOICES is a dataset of natural and spontaneous speech containing a total of 23.7K hours of read (8%), extempore (76%) and conversational (15%) audio from 51K speakers covering 400+ Indian districts and 22 languages. Of these 23.7K hours, 11.2K hours have… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicVoices.audio1M<n<10M115 likes25k downloads3mo agoHugging Face07cfilt /IITB-IndicMonoDocIITB Document level Monolingual Corpora for Indian languages. 22 scheduled languages of India + English (1) Assamese, (2) Bengali, (3) Gujarati, (4) Hindi, (5) Kannada, (6) Kashmiri, (7) Konkani, (8) Malayalam, (9) Manipuri, (10) Marathi, (11) Nepali, (12) Oriya, (13) Punjabi, (14) Sanskrit, (15) Sindhi, (16) Tamil, (17) Telugu, (18) Urdu (19) Bodo, (20) Santhali, (21) Maithili and (22) Dogri. Language Total (#Mil Tokens) bn 5258.47 en 11986.53 gu 887.18 hi 11268.33 kn… See the full description on the dataset page: https://huggingface.co/datasets/cfilt/IITB-IndicMonoDoc.text-generation10B<n<100B11 likes21k downloads2y agoHugging Face08u5753411 /MIT-Indoor-Scenesimagen<1K0 likes16k downloads5mo agoHugging Face09vaquill /open-india-lawgated Open India Law Open, structured Indian primary law - plus the scrapers that build it. Every judgment of the Supreme Court of India and all 25 High Courts, the decisions of 15 tribunals and regulators, and Central, State and Union Territory legislation down to the individual section. Normalized to one schema, exclusively from official government sources. Volume Period Court judgments 12,848,644 1950 to 2025 Tribunal and regulator matters 813,168 1985 to 2026… See the full description on the dataset page: https://huggingface.co/datasets/vaquill/open-india-law.tabulartext-retrieval10M<n<100M26 likes14k downloads1mo agoHugging Face10AlphaDojo /dojo_fin_indicators Languages: 简体中文 · English dojo_fin_indicators — Financial Metrics Overview Multi-period financial statement derivatives per symbol: income statement, balance sheet, cash flow, and industry-specific metrics (banks, insurers, brokers, etc.). Supports quarterly, cumulative, and other report_type values. Files File Description data.parquet Full financial metrics (wide table, 100+ columns) Key Fields (common) Field… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_fin_indicators.tabular100K<n<1M0 likes14k downloads10h agoHugging Face11mratanusarkar /Indian-Laws Dataset Card for Indian Laws This is a comprehensive collection of primary legal documents pertinent to the Indian legal system. It is designed to serve as a foundational resource for supervised fine-tuning (SFT) to make language models, particularly those focused on legal applications tailored for Indian law. text10K<n<100K8 likes12k downloads3y agoHugging Face12binhduong86224 /vineyard-vigor-vegetation-index-archive0 likes12k downloads2d agoHugging Face13inductiva /windtunnel-20k Wind Tunnel Dataset The Wind Tunnel Dataset contains 19,812 OpenFOAM simulations of 1,000 unique automobile-like objects placed in a virtual wind tunnel measuring 20 meters long, 10 meters wide, and 8 meters high. Each object was tested under 20 different conditions: 4 random wind speeds ranging from 10 to 50 m/s, and 5 rotation angles (0°, 180° and 3 random angles). The object meshes were generated using Instant Mesh based on images sourced from the Stanford Cars Dataset. To… See the full description on the dataset page: https://huggingface.co/datasets/inductiva/windtunnel-20k.3dfeature-extraction10K<n<100K7 likes9.9k downloads1y agoHugging Face14ruggsea /infini-news-index INFINI-NEWS FM-Index 🔎 Live search API: these FM-indexes power a public search service — full-text search, n-gram counts, and document retrieval in the browser or via a keyless REST API, without building the index yourself — at infini-news.uni-graz.at (API reference). Pre-built FM-indexes (Burrows–Wheeler Transform + suffix array, built with infini-gram-mini, Liu et al. 2025) over the ruggsea/infini-news-corpus parquets. Enables exact, byte-level substring count and document… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-index.text-retrieval2 likes9.8k downloads8d agoHugging Face15PeterJinGo /wiki-18-e5-index2 likes9.5k downloads2y agoHugging Face16Voxel51 /IndoorSceneRecognition Dataset Card for IndoorSceneRecognition The database contains 67 Indoor categories, and a total of 15620 images. The number of images varies across categories, but there are at least 100 images per category. All images are in jpg format. This is a FiftyOne dataset with 15620 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo import fiftyone.utils.huggingface as fouh # Load the dataset… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/IndoorSceneRecognition.imageimage-classificationn<1K4 likes9.2k downloads2y agoHugging Face17thetrademarkk /india-index-options-1m India Index & Options - 1-minute OHLC 1-minute OHLCV(+OI) bars for NSE/BSE index spot and option chains: NIFTY, BANKNIFTY, SENSEX (~2021-2026). Powers the open-source TradeMarkk backtester (https://thetrademarkk.com). Educational use only. Provided as-is, no warranty. Verify against official exchange data before relying on it. Structure index/{SYMBOL}.parquet - 1-min spot OHLC per index. options/{SYMBOL}/{EXPIRY}.parquet - 1-min OHLC per option contract (with… See the full description on the dataset page: https://huggingface.co/datasets/thetrademarkk/india-index-options-1m.tabular100M<n<1B4 likes8.8k downloads3mo agoHugging Face18wayu-ai /thai-commoncrawl-index Thai Common Crawl Index (2019–2026) An index of every page Common Crawl detected as Thai across 70 monthly crawls, from January 2019 (CC-MAIN-2019-04) to August 2026 (CC-MAIN-2026-30). 932,874,727 page captures · 450,971,497 unique URLs · 6,997,185 hosts · 6,674,969 domains Each row records where the page lives inside Common Crawl's WARC archives — file name, byte offset, and record length — so you can fetch exactly the pages you want with HTTP range requests, without scanning… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-commoncrawl-index.tabular100M<n<1B0 likes8.1k downloads1mo agoHugging Face19ai4bharat /indicvoices_rgated IndicVoices-R: Multilingual, Multi-Speaker Speech Corpus for Indian TTS Dataset Summary IndicVoices-R (IV-R) is the largest multilingual Indian text-to-speech (TTS) dataset derived from an automatic speech recognition (ASR) dataset. It contains 1,704 hours of high-quality speech from 10,496 speakers across 22 Indian languages. This dataset is designed to enhance the development of robust Indian TTS models by providing diverse speaker demographics, natural… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/indicvoices_r.audiotext-to-speech100K<n<1M44 likes7.8k downloads2y agoHugging Face20BAAI /IndustryCorpus[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus.texttext-generation100M<n<1B61 likes7.8k downloads1mo agoHugging Face21tfqdeadlo /Inddatainonefiletext1B<n<10B0 likes7.4k downloads4mo agoHugging Face22omnibioai /pubmed-faiss-indexes0 likes7.3k downloads45m agoHugging Face23commoncrawl /host-index-testing-v2 Common Crawl Host Index v2 GitHub: https://github.com/commoncrawl/cc-host-index Each crawl, we generate a Host Index, which aggregates information about each web hosted visited during the crawl. The information is aggregated from the Common Crawl columnar index, web graph, and raw crawler logs. Quickstart The dataset is Hive-partitioned on crawl (data/crawl=CC-MAIN-2025-18/*.parquet). Open the whole dataset once, then filter with WHERE crawl = '...': because… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/host-index-testing-v2.tabulartext-generation1B<n<10B0 likes7.2k downloads10d agoHugging Face24alibaba-multimodal-industrial-ai /IndustryBench-MIPU IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products Multi-Image Industrial Product Understanding Benchmark — evaluating MLLMs on structured attribute extraction from real-world industrial product images. Industrial product specifications are scattered across multiple heterogeneous images — specification tables, nameplates, technical drawings. IndustryBench-MIPU tests whether MLLMs can reliably recover them through four… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench-MIPU.imageimage-to-text10K<n<100K7 likes6.9k downloads2mo agoHugging Face25jjldo21 /IndustrialDetectionStaticCamerasThe IndustrialDetectionStaticCameras dataset has been collected in order to validate the methodology presented in the paper entitled A few-shot learning methodology for improving safety in industrial scenarios through universal self-supervised visual features and dense optical flow. This dataset is divided into five main folders named videoY, where Y=1,2,3,4,5. Each videoY folder contains the following: The video of the scene in .mp4 format: videoY.mp4 A folder with the images of each frame… See the full description on the dataset page: https://huggingface.co/datasets/jjldo21/IndustrialDetectionStaticCameras.imageobject-detection1K<n<10K1 likes6.1k downloads2y agoHugging Face26castorini /prebuilt-indexes-beir Prebuilt Indexes for BEIR Available indexes: Lucene Flat beir-v1.0.0-trec-covid.bge-base-en-v1.5.flat [readme] Lucene flat index of BEIR collection 'trec-covid' encoded by BGE-base-en-v1.5. beir-v1.0.0-bioasq.bge-base-en-v1.5.flat [readme] Lucene flat index of BEIR collection 'bioasq' encoded by BGE-base-en-v1.5. beir-v1.0.0-nfcorpus.bge-base-en-v1.5.flat [readme] Lucene flat index of BEIR collection 'nfcorpus' encoded by BGE-base-en-v1.5. beir-v1.0.0-nq.bge-base-en-v1.5.flat… See the full description on the dataset page: https://huggingface.co/datasets/castorini/prebuilt-indexes-beir.1 likes5.7k downloads1y agoHugging Face27gagandeepreehal /minuszero-indian-autonomous-driving-dataset-v2gated INDUS-AD: Indian Dataset of Unstructured Urban Scenes for Autonomous Driving Overview INDUS-AD is the largest publicly released Indian autonomous-driving dataset for end-to-end autonomous-driving research. Its name expands to Indian Dataset of Unstructured Urban Scenes for Autonomous Driving. This gated dataset is the decoded companion to the Minus Zero Indian Urban Autonomous Driving Dataset. It provides directly usable camera MP4s, normalized sensor tables… See the full description on the dataset page: https://huggingface.co/datasets/gagandeepreehal/minuszero-indian-autonomous-driving-dataset-v2.imagerobotics10M<n<100M1 likes5.4k downloads7d agoHugging Face28castorini /prebuilt-indexes-msmarco-v1 Prebuilt Indexes for MS MARCO v1 Available indexes: Lucene Standard Inverted msmarco-v1-doc [readme] Lucene index of the MS MARCO V1 document corpus. msmarco-v1-doc-slim [readme] Lucene index of the MS MARCO V1 document corpus ('slim' version). msmarco-v1-doc-full [readme] Lucene index of the MS MARCO V1 document corpus ('full' version). msmarco-v1-doc.d2q-t5 [readme] Lucene index of the MS MARCO V1 document corpus with doc2query-T5 expansions.… See the full description on the dataset page: https://huggingface.co/datasets/castorini/prebuilt-indexes-msmarco-v1.0 likes5.4k downloads8mo agoHugging Face29vdivyasharma /IndicSynth IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research 🏆 Outstanding Paper Award, ACL 2025 🧠 Overview IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/vdivyasharma/IndicSynth.audioaudio-classification1M<n<10M14 likes5.2k downloads9mo agoHugging Face30ManHa /asset-idea-index2 likes5k downloads14h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.