CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Ichsan2895 /alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian We wrangled the original dataset format to 'input' & 'output' format. For example: BEFORE: [ { "from": "human", "value": "Saranlah slogan untuk kampanye daur ulang\n" }, { "from": "gpt", "value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \ "Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \ "Daur… See the full description on the dataset page: https://huggingface.co/datasets/Ichsan2895/alpaca-gpt4-indonesian.textquestion-answering10K<n<100K23 likes861 downloads3y agoHugging Face02l3cube-pune /IndicQuest-v2 IndicQuest v2 A gold-standard multilingual question-answering benchmark for evaluating the India-specific factual knowledge of Large Language Models. 3,471 curriculum-grounded English question–answer pairs across nine domains, translated into 19 Indic languages: 69,420 parallel pairs across 20 languages. More details can be found in our paper. Dataset structure One CSV per language, named <language>.csv (english.csv, hindi.csv, marathi.csv, …). Every file has the… See the full description on the dataset page: https://huggingface.co/datasets/l3cube-pune/IndicQuest-v2.textquestion-answering10K<n<100K0 likes308 downloads1mo agoHugging Face03alibaba-multimodal-industrial-ai /IndustryBench IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs 💻Github | 📝Paper IndustryBench is a multi-lingual benchmark for evaluating the industrial domain knowledge of large language models. It comprises 2,049 expert-curated QA pairs spanning 12 industrial sectors, with human-reviewed translations in Chinese, English, Russian, and Vietnamese. Overview Dimension Details Total questions 2,049 Languages Chinese (zh), English (en), Russian (ru)… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench.textquestion-answering1K<n<10K30 likes204 downloads5mo agoHugging Face04smartduketech /indian-government-schemes-2025 Indian Government Schemes Dataset 2026 Dataset Description The most comprehensive structured dataset of Indian central and state government schemes — 4,693 schemes across all ministries and states, with machine-readable eligibility fields. Maintained by SmartDuke Technologies · Coimbatore, Tamil Nadu, India This dataset powers SchemeFit — India's government scheme finder for citizens and businesses. What Makes This Different Most existing Indian… See the full description on the dataset page: https://huggingface.co/datasets/smartduketech/indian-government-schemes-2025.tabulartext-classification1K<n<10K0 likes190 downloads3mo agoHugging Face05indolem /IndoCareer Introduction IndoCareer is a dataset comprising 8,834 multiple-choice questions designed to evaluate performance in vocational and professional certification exams across various fields. With a focus on Indonesia, IndoCareer provides rich local contexts, spanning six key sectors: (1) healthcare, (2) insurance and finance, (3) creative and design, (4) tourism and hospitality, (5) education and training, and (6) law. Data Each question in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/indolem/IndoCareer.tabularquestion-answering10K<n<100K5 likes121 downloads2y agoHugging Face06cofacto /IndFin-Bench IndFin-Bench: A Benchmark Grounded in Indian Financial Filings Published by Cofacto (formerly CompoundingAI), an AI research platform for Indian equities. IndFin-Bench is a benchmark of 100 hand-curated questions sourced from the corporate filings of Indian listed companies. It is designed to evaluate how accurately LLMs can retrieve and reason over India-specific financial data. Existing financial benchmarks like FinBen and FinQA are built on US market data — SEC filings… See the full description on the dataset page: https://huggingface.co/datasets/cofacto/IndFin-Bench.textquestion-answeringn<1K1 likes97 downloads15d agoHugging Face07Firmansyah-Ibrahim /indo-bloom-corpus 🇮🇩 Indo-Bloom-AQG: A Unified Framework for Controllable Indonesian AQG ⚠️ RESEARCH ARTIFACT STATUS: SILVER VERSION (Work in Progress) This dataset serves as the preliminary corpus (Silver Standard) for the ongoing Doctoral Dissertation at Universitas Negeri Malang (UM). Current State: Unannotated / Pre-validation with Heuristic Bloom Labels Target Final State: Gold Standard (Expert Validated with Bloom's Taxonomy Labels) 🔒 FROZEN — v0.1 Silver This version is permanently… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/indo-bloom-corpus.tabulartext-generation1K<n<10K0 likes83 downloads7mo agoHugging Face08buildwithdmytro /llm-misinformation-resistance-index LLM Misinformation Resistance Index (LMRI) Formal name: LLM Misinformation Resistance Index (LMRI). Public alias: the Gaslighting Index — the two headline scores keep their code names GI-basic and GI-strict, where "GI" comes from the benchmark's public alias. LMRI measures whether a language model will stand up to its own misinformation. Each benchmark item is a fabricated conversation in which the assistant's own prior turn contains a planted false claim (or, for controls, a… See the full description on the dataset page: https://huggingface.co/datasets/buildwithdmytro/llm-misinformation-resistance-index.tabulartext-generation10K<n<100K0 likes54 downloads1mo agoHugging Face09Faishal-Anwar /alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian We wrangled the original dataset format to 'input' & 'output' format. For example: BEFORE: [ { "from": "human", "value": "Saranlah slogan untuk kampanye daur ulang\n" }, { "from": "gpt", "value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \ "Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \ "Daur… See the full description on the dataset page: https://huggingface.co/datasets/Faishal-Anwar/alpaca-gpt4-indonesian.textquestion-answering10K<n<100K0 likes53 downloads20d agoHugging Face10siddharthgaur /indic-reg-bench Dataset Card — indic-reg-bench Status: under construction. The gold set does not exist yet. This card describes what is built, what is not, and the decisions taken so far. It will be wrong in places until labelling is finished; it is published early so the construction method is auditable rather than reconstructed afterwards. What this is A benchmark for Indian regulatory document understanding, built on SEBI (Securities and Exchange Board of India) enforcement… See the full description on the dataset page: https://huggingface.co/datasets/siddharthgaur/indic-reg-bench.texttext-classification10K<n<100K0 likes52 downloads2d agoHugging Face11Pravincoder /Indian_traffic_law_QA Dataset Card for Indian Traffic Rules Dataset Summary This DataSet is curated to Train or fine-tuning LLMs on basic questions on Indian traffic rules. Licensing Information :- bigcode-openrail-m textquestion-answeringn<1K2 likes45 downloads3y agoHugging Face12Ichsan2895 /OASST_Top1_IndonesianBase dataset : OpenAssistant/oasst1 We selected the dataset with "english" language & has rank = 1st. Finally, we translate it to Indonesian with Marian NMT and pretrained model from Helsinki-NLP/opus-mt-en-id. CITATION @InProceedings{mariannmt, title = {Marian: Fast Neural Machine Translation in {C++}}, author = {Junczys-Dowmunt, Marcin and Grundkiewicz, Roman and Dwojak, Tomasz and Hoang, Hieu and Heafield, Kenneth and Neckermann… See the full description on the dataset page: https://huggingface.co/datasets/Ichsan2895/OASST_Top1_Indonesian.textquestion-answering1K<n<10K6 likes40 downloads3y agoHugging Face13glhpradipta /alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian We wrangled the original dataset format to 'input' & 'output' format. For example: BEFORE: [ { "from": "human", "value": "Saranlah slogan untuk kampanye daur ulang\n" }, { "from": "gpt", "value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \ "Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \ "Daur… See the full description on the dataset page: https://huggingface.co/datasets/glhpradipta/alpaca-gpt4-indonesian.textquestion-answering10K<n<100K0 likes40 downloads15d agoHugging Face14Yokey20 /indian-government-schemes-2025 Indian Government Schemes Dataset 2026 Dataset Description The most comprehensive structured dataset of Indian central and state government schemes — 4,693 schemes across all ministries and states, with machine-readable eligibility fields. Maintained by SmartDuke Technologies · Coimbatore, Tamil Nadu, India This dataset powers SchemeFit — India's government scheme finder for citizens and businesses. What Makes This Different Most existing Indian… See the full description on the dataset page: https://huggingface.co/datasets/Yokey20/indian-government-schemes-2025.tabulartext-classification1K<n<10K0 likes39 downloads9d agoHugging Face15jkyung2 /korean-industrial-intelligence-bazaar 🇰🇷 Korea High-Value Industrial Intelligence & Supply-Chain (10,000 Sample Edition) This dataset provides a curated 10,000-record premium showcase of South Korea's high-value industrial supply-chain, market-share, and technological intelligence. ⚡ Need the full 14,300,000+ real-time database?Query our live multi-channel B2A API Gateway directly for 0.01 USDC / USDT per query (Base L2 & Solana):Official Live API: https://husband-voltage-bass-incidents.trycloudflare.com… See the full description on the dataset page: https://huggingface.co/datasets/jkyung2/korean-industrial-intelligence-bazaar.texttext-retrieval10K<n<100K0 likes34 downloads2d agoHugging Face16adalat-ai /Indian-Legal-Retrieval-Generationgated Indian-Legal-Retrieval-Generation An expert-verified evaluation set for retrieval-augmented question answering over Indian court / legal documents. This is the small benchmark used in CourtNav. Paper: CourtNav: Voice-Guided, Anchor-Accurate Navigation of Long Legal Documents in Courtrooms — Sai Khadloya, Kush Juvekar, Arghya Bhattacharya, Utkarsh Saxena. Status: work in progress — contents and structure may still evolve. Overview 21 lawyer-verified question/answer… See the full description on the dataset page: https://huggingface.co/datasets/adalat-ai/Indian-Legal-Retrieval-Generation.documentquestion-answeringn<1K0 likes33 downloads5mo agoHugging Face17Medzza /oncology-financial-reasoning-india 🩺 Medzz-AI: Oncology & Financial Reasoning (India) Status: Active | Context: Indian Healthcare | Focus: Clinical + Economic Logic 👋 The Problem: Why Current Medical AI Fails State-of-the-art LLMs excel at clinical diagnosis but often fail at Health Economics. When asked to generate treatment plans, they frequently hallucinate costs, ignore local insurance constraints, or suggest financially viable treatments that are practically impossible for the patient. Medzz-AI… See the full description on the dataset page: https://huggingface.co/datasets/Medzza/oncology-financial-reasoning-india.textquestion-answeringn<1K0 likes25 downloads9mo agoHugging Face18DevchandraSah /indian-credit-card-facts Indian Credit Card Facts — an open dataset A dated, machine-readable record of the Indian credit-card market: what each card charges, what it earns, how its terms have changed over time, where its points can be transferred, and what those points are worth. Four tables, JSON and CSV, CC BY 4.0. Every release is immutable and separately citable. Table Rows Unit As of cards 280 cards 2026-08-15 changes 1617 events 2026-08-13 transfers 153 edges 2026-07-02… See the full description on the dataset page: https://huggingface.co/datasets/DevchandraSah/indian-credit-card-facts.texttabular-classification1K<n<10K0 likes24 downloads1mo agoHugging Face19edyfjm07 /squad_indicaciones_estextquestion-answering1K<n<10K1 likes23 downloads2y agoHugging Face20Firmansyah-Ibrahim /IndoBloom-AQG-Benchmark-Corpus 📚 Indo-Bloom AQG Benchmark Corpus (All Models) ⚠️ RESEARCH ARTIFACT STATUS: BENCHMARK / SILVER CORPUS (Stage 1) This dataset serves as the comparative benchmark corpus for the Indo-Bloom research project at Universitas Negeri Malang (UM). Current State: LLM-Generated QA Pairs Evaluated via Rule-Based Evaluator Next Stage: Expert Annotation (Stage 2) → Gold Standard 🔒 FROZEN — Benchmark v1.0 This version is permanently frozen to ensure reproducibility of the experimental… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/IndoBloom-AQG-Benchmark-Corpus.tabulartext-generation10K<n<100K0 likes23 downloads6mo agoHugging Face21anon-neurips26-3390 /IndustryBench IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs Anonymized review copy — NeurIPS 2026 Evaluations & Datasets Track, Submission 3390 (under review). IndustryBench is a benchmark for evaluating the industrial procurement knowledge of large language models. It comprises 2,049 QA pairs grounded in Chinese national standards (GB/T) and structured industrial product records, with item-aligned renderings in Chinese, English, Russian, and Vietnamese. Erratum note:… See the full description on the dataset page: https://huggingface.co/datasets/anon-neurips26-3390/IndustryBench.textquestion-answering1K<n<10K0 likes22 downloads2mo agoHugging Face22Alwiiiiiiiiii /indo-bloom-raw-bse 📚 Indo-Bloom BSE RAW Corpus ⚠️ RESEARCH ARTIFACT STATUS: RAW CORPUS (Stage 0) This dataset serves as the raw material corpus for the Indo-Bloom research project at Universitas Negeri Malang (UM). Current State: Extracted & Cleaned Context from BSE Textbooks Next Stage: QA Pair Generation (Stage 1) → Silver Corpus 🔒 FROZEN — Raw v1.0 This version is permanently frozen to ensure reproducibility. This corpus will be used as input for QA generation pipeline. 📄… See the full description on the dataset page: https://huggingface.co/datasets/Alwiiiiiiiiii/indo-bloom-raw-bse.tabulartext-generation1K<n<10K0 likes22 downloads3d agoHugging Face23Firmansyah-Ibrahim /indo-bloom-raw-bse 📚 Indo-Bloom BSE RAW Corpus ⚠️ RESEARCH ARTIFACT STATUS: RAW CORPUS (Stage 0) This dataset serves as the raw material corpus for the Indo-Bloom research project at Universitas Negeri Malang (UM). Current State: Extracted & Cleaned Context from BSE Textbooks Next Stage: QA Pair Generation (Stage 1) → Silver Corpus 🔒 FROZEN — Raw v1.0 This version is permanently frozen to ensure reproducibility. This corpus will be used as input for QA generation pipeline. 📄 Associated… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/indo-bloom-raw-bse.tabulartext-generation1K<n<10K0 likes21 downloads7mo agoHugging Face24siva0072 /indian-government-schemes-2025 Indian Government Schemes Dataset 2026 Dataset Description The most comprehensive structured dataset of Indian central and state government schemes — 4,693 schemes across all ministries and states, with machine-readable eligibility fields. Maintained by SmartDuke Technologies · Coimbatore, Tamil Nadu, India This dataset powers SchemeFit — India's government scheme finder for citizens and businesses. What Makes This Different Most existing Indian… See the full description on the dataset page: https://huggingface.co/datasets/siva0072/indian-government-schemes-2025.tabulartext-classification1K<n<10K0 likes20 downloads2mo agoHugging Face25Sumna /indian-government-schemes-2025gated Indian Government Schemes Dataset 2026 Dataset Description The most comprehensive structured dataset of Indian central and state government schemes — 4,693 schemes across all ministries and states, with machine-readable eligibility fields. Maintained by SmartDuke Technologies · Coimbatore, Tamil Nadu, India This dataset powers SchemeFit — India's government scheme finder for citizens and businesses. What Makes This Different Most existing Indian… See the full description on the dataset page: https://huggingface.co/datasets/Sumna/indian-government-schemes-2025.tabulartext-classification1K<n<10K0 likes20 downloads7d agoHugging Face26steamed-potatop /indian-family-law-data Indian Family Law QA Dataset This dataset contains 1,401 high-quality question-answer pairs focused on Family Law II, specifically tailored for legal studies in the Indian context. The dataset covers critical areas such as Hindu Succession, Coparcenary rights, and principles of Muslim Law. Dataset Details Total Rows: 1,401 Language: English Task: Question Answering / Legal Knowledge Retrieval Format: CSV (two-column format) Dataset Structure The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/steamed-potatop/indian-family-law-data.textquestion-answering1K<n<10K0 likes19 downloads9mo agoHugging Face27AmalAkbar /alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian We wrangled the original dataset format to 'input' & 'output' format. For example: BEFORE: [ { "from": "human", "value": "Saranlah slogan untuk kampanye daur ulang\n" }, { "from": "gpt", "value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \ "Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \ "Daur… See the full description on the dataset page: https://huggingface.co/datasets/AmalAkbar/alpaca-gpt4-indonesian.textquestion-answering10K<n<100K0 likes13 downloads2mo agoHugging Face28akahana /alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian We wrangled the original dataset format to 'input' & 'output' format. For example: BEFORE: [ { "from": "human", "value": "Saranlah slogan untuk kampanye daur ulang\n" }, { "from": "gpt", "value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \ "Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \ "Daur… See the full description on the dataset page: https://huggingface.co/datasets/akahana/alpaca-gpt4-indonesian.textquestion-answering10K<n<100K1 likes10 downloads9mo agoHugging Face29har123ish /indian-government-schemes-2025 Indian Government Schemes Dataset 2026 Dataset Description The most comprehensive structured dataset of Indian central and state government schemes — 4,693 schemes across all ministries and states, with machine-readable eligibility fields. Maintained by SmartDuke Technologies · Coimbatore, Tamil Nadu, India This dataset powers SchemeFit — India's government scheme finder for citizens and businesses. What Makes This Different Most existing Indian… See the full description on the dataset page: https://huggingface.co/datasets/har123ish/indian-government-schemes-2025.tabulartext-classification1K<n<10K0 likes10 downloads1mo agoHugging Face30cyberblip /india2tabularquestion-answering1K<n<10K0 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.