CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01vaquill /open-india-lawgated Open India Law Open, structured Indian primary law - plus the scrapers that build it. Every judgment of the Supreme Court of India and all 25 High Courts, the decisions of 15 tribunals and regulators, and Central, State and Union Territory legislation down to the individual section. Normalized to one schema, exclusively from official government sources. Volume Period Court judgments 12,848,644 1950 to 2025 Tribunal and regulator matters 813,168 1985 to 2026… See the full description on the dataset page: https://huggingface.co/datasets/vaquill/open-india-law.tabulartext-retrieval10M<n<100M26 likes14k downloads1mo agoHugging Face02hemanthreddy901 /nadora-global-industries NADORA Global Industries A synthetic multinational, built to be developed against rather than demonstrated with. One fictional company — $5.20bn revenue, $716m EBITDA, 24,000 employees, 18 countries, 35 legal entities, five business units — traded daily from January 2022 to December 2026 and rendered at six fidelities, from a 3 MB unit-test fixture to a 5 GB full-scale corpus. 38,964,663 rows · 11 GB · 2,319 verification assertions, all passing. 100% synthetic. No real company… See the full description on the dataset page: https://huggingface.co/datasets/hemanthreddy901/nadora-global-industries.documenttabular-regression10K<n<100K0 likes724 downloads10d agoHugging Face03open-index /open-library Open Library The complete Open Library catalog in clean, analysis-ready Parquet. 150.0M+ records across 11 entity types, from ISBNs and author bios to reading logs and Wikidata links. What is it? Open Library is a complete snapshot of the Open Library database, an open project of the Internet Archive with the mission of creating "one web page for every book ever published." The catalog is community-edited and contains bibliographic records for millions of authors, works… See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-library.tabulartext-generation100M<n<1B9 likes611 downloads6mo agoHugging Face04daruokta /t5gemma2-indonesia-instruct-v1 T5Gemma-2 Indonesian Instruct — Mono-Repo Satu repositori dataset HF untuk seluruh data pelatihan T5-Gemma-2 bahasa Indonesia. Diorganisasi per fungsi (fondasi → spesifik → preferensi) dengan folder/subfolder, setiap config = folder dan berisi split train + validation (80:20) di level percakapan. Struktur (by fungsi) t5gemma2-indonesia-instruct-v1/ ├── README.md ├── manifest.json ├── chat_idx_map.json ├── foundation/ ← FASE 1 · fondasi Bahasa… See the full description on the dataset page: https://huggingface.co/datasets/daruokta/t5gemma2-indonesia-instruct-v1.imagetext-generation100K<n<1M0 likes519 downloads3d agoHugging Face05AgentCrush /agents-index AgentCrush Agent Index Evidence-ranked index of the AI agent economy. Updated daily from agentcrush.xyz. Overview 1,445 agents indexed across categories: developer tools, tokenized agents, service agents, model families 207 evidence-ranked with verified multi-signal scores Updated: 2026-09-26 Configs Config Description Rows agents All indexed agents with metadata ~1,445 evidence_ranked Evidence-ranked tier only ~207 snapshots_latest… See the full description on the dataset page: https://huggingface.co/datasets/AgentCrush/agents-index.tabulartext-classification1K<n<10K0 likes508 downloads12h agoHugging Face06KanoonGPT /indian-legal-documents Indian Legal Documents Open Indian statutory and legal-document data for AI, search, and legal research. This dataset is part of the KanoonGPT Open Legal Data Initiative — an effort to make Indian legal data easier to access, structure, and build on for open-source research, legal tech, and production AI systems. KanoonGPT is building structured Indian legal datasets and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.… See the full description on the dataset page: https://huggingface.co/datasets/KanoonGPT/indian-legal-documents.texttext-generation10K<n<100K0 likes392 downloads4mo agoHugging Face07BAAI /IndustryInstruction_Finance-Economics IndustryInstruction: Finance & Economics This repository contains the IndustryInstruction: Finance & Economics domain subset of BAAI/IndustryInstruction. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryInstruction: @misc{shi2024industryinstruction, title = {IndustryInstruction}, author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Finance-Economics.tabularquestion-answering100K<n<1M10 likes382 downloads1mo agoHugging Face08Parssky /industrial-instruction-dataset Industrial-Instruction Dataset Industrial-Instruction provides benchmark and training-ready QA instances derived from industrial technical reports, designed to evaluate robustness under realistic retrieval conditions. Samples are grounded in retrieved evidence and include irrelevant retrieval, single-/multi-document support, and single-/multi-document answer settings. Paper Industrial-Instruction: An End-to-End Framework for Building Instruction-Tuning and… See the full description on the dataset page: https://huggingface.co/datasets/Parssky/industrial-instruction-dataset.tabularquestion-answering10K<n<100K0 likes382 downloads1mo agoHugging Face09emperor-mew /global-censorship-index Voidly Global Censorship Index Real-time internet censorship measurements for 200 countries, based on 38,780,449+ OONI network probes. Dataset Description The Global Censorship Index provides country-level internet censorship scores derived from actual network measurements. Unlike annual expert assessments, this data updates daily. Key Statistics Countries covered: 200 Total measurements: 38,780,449 Severe censorship: 1 countries High censorship: 6… See the full description on the dataset page: https://huggingface.co/datasets/emperor-mew/global-censorship-index.tabulartext-classificationn<1K1 likes342 downloads2mo agoHugging Face10LeSouv /index-souverainete Index Souveraineté Le Souv Le dataset public de référence sur les entreprises françaises stratégiques cédées à des capitaux étrangers, et sur les entreprises souveraines à capitaux français. Source canonique : Le Souv — média indépendant consacré à la souveraineté économique et politique française. URL canonique : https://lesouv.fr/index-souverainete.json Licence : Creative Commons BY 4.0 (réutilisation libre avec attribution à Le Souv). Mise à jour : continue, regénération… See the full description on the dataset page: https://huggingface.co/datasets/LeSouv/index-souverainete.tabulartext-classificationn<1K1 likes335 downloads3d agoHugging Face11BAAI /IndustryInstruction_Aerospace IndustryInstruction: Aerospace This repository contains the IndustryInstruction: Aerospace domain subset of BAAI/IndustryInstruction. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryInstruction: @misc{shi2024industryinstruction, title = {IndustryInstruction}, author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao and… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Aerospace.tabularquestion-answering100K<n<1M6 likes322 downloads1mo agoHugging Face12BAAI /IndustryInstruction_Artificial-Intelligence IndustryInstruction: Artificial Intelligence This repository contains the IndustryInstruction: Artificial Intelligence domain subset of BAAI/IndustryInstruction. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryInstruction: @misc{shi2024industryinstruction, title = {IndustryInstruction}, author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Artificial-Intelligence.tabularquestion-answering100K<n<1M2 likes273 downloads1mo agoHugging Face13open-index /luatdo-graph luatdo-graph A knowledge graph over Vietnamese law, built from about 128,000 documents by luatdo. This is the result of running the pipeline, published so that nobody has to run it again. The pipeline takes days and several hundred dollars of model calls, and the output is the same for everyone. What is in it Nodes 8,175,346 Relationships 9,119,011 Node tables 14 Relationship tables 19 Node labels 13 Parquet 623MB across 47 files Neo4j… See the full description on the dataset page: https://huggingface.co/datasets/open-index/luatdo-graph.tabularquestion-answering10M<n<100M0 likes265 downloads2mo agoHugging Face14BAAI /IndustryInstruction_Technology-Research IndustryInstruction: Technology & Research This repository contains the IndustryInstruction: Technology & Research domain subset of BAAI/IndustryInstruction. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryInstruction: @misc{shi2024industryinstruction, title = {IndustryInstruction}, author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Technology-Research.tabularquestion-answering100K<n<1M0 likes217 downloads1mo agoHugging Face15BAAI /IndustryInstruction_Automobiles IndustryInstruction: Automobiles This repository contains the IndustryInstruction: Automobiles domain subset of BAAI/IndustryInstruction. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryInstruction: @misc{shi2024industryinstruction, title = {IndustryInstruction}, author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Automobiles.tabularquestion-answering100K<n<1M5 likes215 downloads1mo agoHugging Face16BAAI /IndustryInstruction_Transportation IndustryInstruction: Transportation This repository contains the IndustryInstruction: Transportation domain subset of BAAI/IndustryInstruction. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryInstruction: @misc{shi2024industryinstruction, title = {IndustryInstruction}, author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Transportation.tabularquestion-answering100K<n<1M1 likes210 downloads1mo agoHugging Face17BAAI /IndustryInstruction_Literature-Emotions IndustryInstruction: Literature & Emotions This repository contains the IndustryInstruction: Literature & Emotions domain subset of BAAI/IndustryInstruction. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryInstruction: @misc{shi2024industryinstruction, title = {IndustryInstruction}, author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Literature-Emotions.tabularquestion-answering100K<n<1M2 likes189 downloads1mo agoHugging Face18BAAI /IndustryInstruction_Law-Justice IndustryInstruction: Law & Justice This repository contains the IndustryInstruction: Law & Justice domain subset of BAAI/IndustryInstruction. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryInstruction: @misc{shi2024industryinstruction, title = {IndustryInstruction}, author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Law-Justice.tabularquestion-answering100K<n<1M1 likes186 downloads1mo agoHugging Face19smartduketech /indian-government-schemes-2025 Indian Government Schemes Dataset 2026 Dataset Description The most comprehensive structured dataset of Indian central and state government schemes — 4,693 schemes across all ministries and states, with machine-readable eligibility fields. Maintained by SmartDuke Technologies · Coimbatore, Tamil Nadu, India This dataset powers SchemeFit — India's government scheme finder for citizens and businesses. What Makes This Different Most existing Indian… See the full description on the dataset page: https://huggingface.co/datasets/smartduketech/indian-government-schemes-2025.tabulartext-classification1K<n<10K0 likes176 downloads3mo agoHugging Face20BAAI /IndustryInstruction_Travel-Geography IndustryInstruction: Travel & Geography This repository contains the IndustryInstruction: Travel & Geography domain subset of BAAI/IndustryInstruction. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryInstruction: @misc{shi2024industryinstruction, title = {IndustryInstruction}, author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Travel-Geography.tabularquestion-answering100K<n<1M0 likes157 downloads1mo agoHugging Face21BAAI /IndustryInstruction_Hospitality-Catering IndustryInstruction: Hospitality Catering This repository contains the IndustryInstruction: Hospitality Catering domain subset of BAAI/IndustryInstruction. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryInstruction: @misc{shi2024industryinstruction, title = {IndustryInstruction}, author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Hospitality-Catering.tabularquestion-answering100K<n<1M1 likes150 downloads1mo agoHugging Face22BAAI /IndustryInstructiongated本数据集为行业指令数据集,目前包含的行业中英文对照名称如下,本次数据旨在补充当前行业指令数据的空白,并挖掘BAAI/IndustryCorpus2预训练数据集中高质量预训练语料中包含的行业高价值知识。 汽车 : Automobiles 航空航天 : Aerospace 人工智能_机器学习 : Artificial-Intelligence 交通运输 : Transportation 科技_科学研究 : Technology-Research 法律_司法 : Law-Justice 金融_经济 : Finance-Economics 文学_情感 : Literature-Emotions 旅游_地理 : Travel-Geography 住宿_餐饮_酒店 : Hospitality-Catering 医疗 : Health-Medicine 学科教育 : Subject-Education 我们为每个数据集目录下面都提供了对应行业数据的 词云可视化和 数据质量分布曲线。如果需要单独行业的数据,可以跳转到单独的行业数据集地址… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction.imagequestion-answering1M<n<10M33 likes136 downloads1mo agoHugging Face23roshankaranth /IndicDB IndicDB — Multilingual Text-to-SQL Benchmark for Indian Languages IndicDB is a comprehensive multilingual Text-to-SQL benchmark for evaluating cross-lingual semantic parsing across diverse Indic language families. Questions are posed in 7 languages while the underlying schemas and values remain in English — simultaneously stressing translation, schema linking, value grounding, and multi-table join reasoning. Schemas are sourced from real Indian open-data platforms (NDAP and… See the full description on the dataset page: https://huggingface.co/datasets/roshankaranth/IndicDB.tabulartable-question-answering10K<n<100K0 likes134 downloads3mo agoHugging Face24indolem /IndoCareer Introduction IndoCareer is a dataset comprising 8,834 multiple-choice questions designed to evaluate performance in vocational and professional certification exams across various fields. With a focus on Indonesia, IndoCareer provides rich local contexts, spanning six key sectors: (1) healthcare, (2) insurance and finance, (3) creative and design, (4) tourism and hospitality, (5) education and training, and (6) law. Data Each question in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/indolem/IndoCareer.tabularquestion-answering10K<n<100K5 likes120 downloads2y agoHugging Face25bio-protocol /neophyte-faiss-index-v1 neophyte-faiss-index-v1 A FAISS index + metadata for scientific retrieval Contents index.faiss: FAISS index (cosine w/ inner product). meta.jsonl: one JSON per chunk; fields include chunk_id, paper_id, title, section, subsection, paragraph_index, keywords, boost. index.info.json: (optional) dimensions, index type, faiss version. Build provenance Chunking: hierarchical (section→paragraph→~480-token chunks, ~15% overlap) Embedder:… See the full description on the dataset page: https://huggingface.co/datasets/bio-protocol/neophyte-faiss-index-v1.tabulartext-retrieval100K<n<1M0 likes89 downloads11mo agoHugging Face26jakartaresearch /indoqaThis dataset is built for question answering task.tabularquestion-answering1K<n<10K13 likes87 downloads4y agoHugging Face27ibahasa /indommlu-local-languages IndoMMLU: Local Languages and Cultures (audited subset) An audited, corrected subset of IndoMMLU (Koto et al., 2023) covering the 9 Local Languages and Cultures subjects: Indonesian primary and secondary school exam questions written in Balinese, Banjarese, Dayak Ngaju, Javanese, Lampung, Madurese, Makassarese, and Sundanese, plus one culture-knowledge subject on Minangkabau customs (answered in standard Indonesian). This is not a dataset we created. It is IndoMMLU's own subset… See the full description on the dataset page: https://huggingface.co/datasets/ibahasa/indommlu-local-languages.tabularmultiple-choice1K<n<10K0 likes87 downloads1mo agoHugging Face28Firmansyah-Ibrahim /indo-bloom-corpus 🇮🇩 Indo-Bloom-AQG: A Unified Framework for Controllable Indonesian AQG ⚠️ RESEARCH ARTIFACT STATUS: SILVER VERSION (Work in Progress) This dataset serves as the preliminary corpus (Silver Standard) for the ongoing Doctoral Dissertation at Universitas Negeri Malang (UM). Current State: Unannotated / Pre-validation with Heuristic Bloom Labels Target Final State: Gold Standard (Expert Validated with Bloom's Taxonomy Labels) 🔒 FROZEN — v0.1 Silver This version is permanently… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/indo-bloom-corpus.tabulartext-generation1K<n<10K0 likes85 downloads7mo agoHugging Face29daruokta /t5gemma2-indonesia-chat-formatted T5Gemma-2 Indonesian Chat & QA Dataset A high-quality Indonesian language multi-turn conversation and reading comprehension dataset, specifically formatted for instruction tuning of sequence-to-sequence (Seq2Seq) models like T5-Gemma / T5-Gemma-2. Dataset Description This dataset contains over 7,400 multi-turn conversations and document-based Q&A in Bahasa Indonesia. It covers diverse topics including everyday life, technology, general knowledge, and structured… See the full description on the dataset page: https://huggingface.co/datasets/daruokta/t5gemma2-indonesia-chat-formatted.tabulartext-generation10K<n<100K1 likes77 downloads3mo agoHugging Face30Tassy24 /K-Paths-inductive-reasoning-drugbank 🔗 This dataset is part of the study: K-Paths: Reasoning over Graph Paths for Drug Repurposing and Drug Interaction Prediction 📖 Read the Paper 💾 GitHub Repository DrugBank: Inductive Reasoning Dataset This dataset contains drug pairs annotated with 86 pharmacological relationships (e.g.,DrugA may increase the anticholinergic activities of DrugB). Each entry includes two drugs, an interaction label, drug descriptions, and structured/natural language representations… See the full description on the dataset page: https://huggingface.co/datasets/Tassy24/K-Paths-inductive-reasoning-drugbank.tabularquestion-answering100K<n<1M0 likes68 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.