datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian
We wrangled the original dataset format to 'input' & 'output' format. For example:
BEFORE:
[ { "from": "human",
"value": "Saranlah slogan untuk kampanye daur ulang\n" },
{ "from": "gpt",
"value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \
"Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \
"Daur… See the full description on the dataset page: https://huggingface.co/datasets/Ichsan2895/alpaca-gpt4-indonesian.IndicQuest-v2
IndicQuest v2
A gold-standard multilingual question-answering benchmark for evaluating the India-specific factual knowledge of Large Language Models. 3,471 curriculum-grounded English question–answer pairs across nine domains, translated into 19 Indic languages: 69,420 parallel pairs across 20 languages.
More details can be found in our paper.
Dataset structure
One CSV per language, named <language>.csv (english.csv, hindi.csv, marathi.csv, …). Every file has the… See the full description on the dataset page: https://huggingface.co/datasets/l3cube-pune/IndicQuest-v2.IndustryBench
IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs
💻Github | 📝Paper
IndustryBench is a multi-lingual benchmark for evaluating the industrial domain knowledge of large language models. It comprises 2,049 expert-curated QA pairs spanning 12 industrial sectors, with human-reviewed translations in Chinese, English, Russian, and Vietnamese.
Overview
Dimension
Details
Total questions
2,049
Languages
Chinese (zh), English (en), Russian (ru)… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench.indian-government-schemes-2025
Indian Government Schemes Dataset 2026
Dataset Description
The most comprehensive structured dataset of Indian central and state government schemes — 4,693 schemes across all ministries and states, with machine-readable eligibility fields.
Maintained by SmartDuke Technologies · Coimbatore, Tamil Nadu, India
This dataset powers SchemeFit — India's government scheme finder for citizens and businesses.
What Makes This Different
Most existing Indian… See the full description on the dataset page: https://huggingface.co/datasets/smartduketech/indian-government-schemes-2025.IndoCareer
Introduction
IndoCareer is a dataset comprising 8,834 multiple-choice questions designed to evaluate performance in vocational and professional certification exams across various fields. With a focus on Indonesia, IndoCareer provides rich local contexts, spanning six key sectors: (1) healthcare, (2) insurance and finance, (3) creative and design, (4) tourism and hospitality, (5) education and training, and (6) law.
Data
Each question in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/indolem/IndoCareer.IndFin-Bench
IndFin-Bench: A Benchmark Grounded in Indian Financial Filings
Published by Cofacto (formerly CompoundingAI), an AI research platform for Indian equities.
IndFin-Bench is a benchmark of 100 hand-curated questions sourced from the corporate filings of Indian listed companies. It is designed to evaluate how accurately LLMs can retrieve and reason over India-specific financial data.
Existing financial benchmarks like FinBen and FinQA are built on US market data — SEC filings… See the full description on the dataset page: https://huggingface.co/datasets/cofacto/IndFin-Bench.indo-bloom-corpus
🇮🇩 Indo-Bloom-AQG: A Unified Framework for Controllable Indonesian AQG
⚠️ RESEARCH ARTIFACT STATUS: SILVER VERSION (Work in Progress)
This dataset serves as the preliminary corpus (Silver Standard) for the ongoing Doctoral Dissertation at Universitas Negeri Malang (UM).
Current State: Unannotated / Pre-validation with Heuristic Bloom Labels
Target Final State: Gold Standard (Expert Validated with Bloom's Taxonomy Labels)
🔒 FROZEN — v0.1 Silver
This version is permanently… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/indo-bloom-corpus.llm-misinformation-resistance-index
LLM Misinformation Resistance Index (LMRI)
Formal name: LLM Misinformation Resistance Index (LMRI).
Public alias: the Gaslighting Index — the two headline scores keep their code
names GI-basic and GI-strict, where "GI" comes from the benchmark's public alias.
LMRI measures whether a language model will stand up to its own misinformation.
Each benchmark item is a fabricated conversation in which the assistant's own prior
turn contains a planted false claim (or, for controls, a… See the full description on the dataset page: https://huggingface.co/datasets/buildwithdmytro/llm-misinformation-resistance-index.alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian
We wrangled the original dataset format to 'input' & 'output' format. For example:
BEFORE:
[ { "from": "human",
"value": "Saranlah slogan untuk kampanye daur ulang\n" },
{ "from": "gpt",
"value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \
"Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \
"Daur… See the full description on the dataset page: https://huggingface.co/datasets/Faishal-Anwar/alpaca-gpt4-indonesian.indic-reg-bench
Dataset Card — indic-reg-bench
Status: under construction. The gold set does not exist yet. This card describes what is built, what is not, and the decisions taken so far. It will be wrong in places until labelling is finished; it is published early so the construction method is auditable rather than reconstructed afterwards.
What this is
A benchmark for Indian regulatory document understanding, built on SEBI (Securities and Exchange Board of India) enforcement… See the full description on the dataset page: https://huggingface.co/datasets/siddharthgaur/indic-reg-bench.Indian_traffic_law_QA
Dataset Card for Indian Traffic Rules
Dataset Summary
This DataSet is curated to Train or fine-tuning LLMs on basic questions on Indian traffic rules.
Licensing Information :- bigcode-openrail-m
OASST_Top1_IndonesianBase dataset : OpenAssistant/oasst1
We selected the dataset with "english" language & has rank = 1st.
Finally, we translate it to Indonesian with Marian NMT and pretrained model from Helsinki-NLP/opus-mt-en-id.
CITATION
@InProceedings{mariannmt,
title = {Marian: Fast Neural Machine Translation in {C++}},
author = {Junczys-Dowmunt, Marcin and Grundkiewicz, Roman and
Dwojak, Tomasz and Hoang, Hieu and Heafield, Kenneth and
Neckermann… See the full description on the dataset page: https://huggingface.co/datasets/Ichsan2895/OASST_Top1_Indonesian.alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian
We wrangled the original dataset format to 'input' & 'output' format. For example:
BEFORE:
[ { "from": "human",
"value": "Saranlah slogan untuk kampanye daur ulang\n" },
{ "from": "gpt",
"value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \
"Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \
"Daur… See the full description on the dataset page: https://huggingface.co/datasets/glhpradipta/alpaca-gpt4-indonesian.indian-government-schemes-2025
Indian Government Schemes Dataset 2026
Dataset Description
The most comprehensive structured dataset of Indian central and state government schemes — 4,693 schemes across all ministries and states, with machine-readable eligibility fields.
Maintained by SmartDuke Technologies · Coimbatore, Tamil Nadu, India
This dataset powers SchemeFit — India's government scheme finder for citizens and businesses.
What Makes This Different
Most existing Indian… See the full description on the dataset page: https://huggingface.co/datasets/Yokey20/indian-government-schemes-2025.korean-industrial-intelligence-bazaar
🇰🇷 Korea High-Value Industrial Intelligence & Supply-Chain (10,000 Sample Edition)
This dataset provides a curated 10,000-record premium showcase of South Korea's high-value industrial supply-chain, market-share, and technological intelligence.
⚡ Need the full 14,300,000+ real-time database?Query our live multi-channel B2A API Gateway directly for 0.01 USDC / USDT per query (Base L2 & Solana):Official Live API: https://husband-voltage-bass-incidents.trycloudflare.com… See the full description on the dataset page: https://huggingface.co/datasets/jkyung2/korean-industrial-intelligence-bazaar.Indian-Legal-Retrieval-Generation
Indian-Legal-Retrieval-Generation
An expert-verified evaluation set for retrieval-augmented question answering over Indian
court / legal documents. This is the small benchmark used in CourtNav.
Paper: CourtNav: Voice-Guided, Anchor-Accurate Navigation of Long Legal Documents in Courtrooms — Sai Khadloya, Kush Juvekar, Arghya Bhattacharya, Utkarsh Saxena.
Status: work in progress — contents and structure may still evolve.
Overview
21 lawyer-verified question/answer… See the full description on the dataset page: https://huggingface.co/datasets/adalat-ai/Indian-Legal-Retrieval-Generation.oncology-financial-reasoning-india
🩺 Medzz-AI: Oncology & Financial Reasoning (India)
Status: Active | Context: Indian Healthcare | Focus: Clinical + Economic Logic
👋 The Problem: Why Current Medical AI Fails
State-of-the-art LLMs excel at clinical diagnosis but often fail at Health Economics. When asked to generate treatment plans, they frequently hallucinate costs, ignore local insurance constraints, or suggest financially viable treatments that are practically impossible for the patient.
Medzz-AI… See the full description on the dataset page: https://huggingface.co/datasets/Medzza/oncology-financial-reasoning-india.indian-credit-card-facts
Indian Credit Card Facts — an open dataset
A dated, machine-readable record of the Indian credit-card market: what each card charges, what
it earns, how its terms have changed over time, where its points can be transferred, and what
those points are worth.
Four tables, JSON and CSV, CC BY 4.0. Every release is immutable and separately citable.
Table
Rows
Unit
As of
cards
280
cards
2026-08-15
changes
1617
events
2026-08-13
transfers
153
edges
2026-07-02… See the full description on the dataset page: https://huggingface.co/datasets/DevchandraSah/indian-credit-card-facts.squad_indicaciones_esIndoBloom-AQG-Benchmark-Corpus
📚 Indo-Bloom AQG Benchmark Corpus (All Models)
⚠️ RESEARCH ARTIFACT STATUS: BENCHMARK / SILVER CORPUS (Stage 1)
This dataset serves as the comparative benchmark corpus for the Indo-Bloom research project at Universitas Negeri Malang (UM).
Current State: LLM-Generated QA Pairs Evaluated via Rule-Based Evaluator
Next Stage: Expert Annotation (Stage 2) → Gold Standard
🔒 FROZEN — Benchmark v1.0
This version is permanently frozen to ensure reproducibility of the experimental… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/IndoBloom-AQG-Benchmark-Corpus.IndustryBench
IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs
Anonymized review copy — NeurIPS 2026 Evaluations & Datasets Track, Submission 3390 (under review).
IndustryBench is a benchmark for evaluating the industrial procurement knowledge of large language models. It comprises 2,049 QA pairs grounded in Chinese national standards (GB/T) and structured industrial product records, with item-aligned renderings in Chinese, English, Russian, and Vietnamese.
Erratum note:… See the full description on the dataset page: https://huggingface.co/datasets/anon-neurips26-3390/IndustryBench.indo-bloom-raw-bse
📚 Indo-Bloom BSE RAW Corpus
⚠️ RESEARCH ARTIFACT STATUS: RAW CORPUS (Stage 0)
This dataset serves as the raw material corpus for the Indo-Bloom research project at Universitas Negeri Malang (UM).
Current State: Extracted & Cleaned Context from BSE Textbooks
Next Stage: QA Pair Generation (Stage 1) → Silver Corpus
🔒 FROZEN — Raw v1.0
This version is permanently frozen to ensure reproducibility. This corpus will be used as input for QA generation pipeline.
📄… See the full description on the dataset page: https://huggingface.co/datasets/Alwiiiiiiiiii/indo-bloom-raw-bse.indo-bloom-raw-bse
📚 Indo-Bloom BSE RAW Corpus
⚠️ RESEARCH ARTIFACT STATUS: RAW CORPUS (Stage 0)
This dataset serves as the raw material corpus for the Indo-Bloom research project at Universitas Negeri Malang (UM).
Current State: Extracted & Cleaned Context from BSE Textbooks
Next Stage: QA Pair Generation (Stage 1) → Silver Corpus
🔒 FROZEN — Raw v1.0
This version is permanently frozen to ensure reproducibility. This corpus will be used as input for QA generation pipeline.
📄 Associated… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/indo-bloom-raw-bse.indian-government-schemes-2025
Indian Government Schemes Dataset 2026
Dataset Description
The most comprehensive structured dataset of Indian central and state government schemes — 4,693 schemes across all ministries and states, with machine-readable eligibility fields.
Maintained by SmartDuke Technologies · Coimbatore, Tamil Nadu, India
This dataset powers SchemeFit — India's government scheme finder for citizens and businesses.
What Makes This Different
Most existing Indian… See the full description on the dataset page: https://huggingface.co/datasets/siva0072/indian-government-schemes-2025.indian-government-schemes-2025
Indian Government Schemes Dataset 2026
Dataset Description
The most comprehensive structured dataset of Indian central and state government schemes — 4,693 schemes across all ministries and states, with machine-readable eligibility fields.
Maintained by SmartDuke Technologies · Coimbatore, Tamil Nadu, India
This dataset powers SchemeFit — India's government scheme finder for citizens and businesses.
What Makes This Different
Most existing Indian… See the full description on the dataset page: https://huggingface.co/datasets/Sumna/indian-government-schemes-2025.indian-family-law-data
Indian Family Law QA Dataset
This dataset contains 1,401 high-quality question-answer pairs focused on Family Law II, specifically tailored for legal studies in the Indian context. The dataset covers critical areas such as Hindu Succession, Coparcenary rights, and principles of Muslim Law.
Dataset Details
Total Rows: 1,401
Language: English
Task: Question Answering / Legal Knowledge Retrieval
Format: CSV (two-column format)
Dataset Structure
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/steamed-potatop/indian-family-law-data.alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian
We wrangled the original dataset format to 'input' & 'output' format. For example:
BEFORE:
[ { "from": "human",
"value": "Saranlah slogan untuk kampanye daur ulang\n" },
{ "from": "gpt",
"value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \
"Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \
"Daur… See the full description on the dataset page: https://huggingface.co/datasets/AmalAkbar/alpaca-gpt4-indonesian.alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian
We wrangled the original dataset format to 'input' & 'output' format. For example:
BEFORE:
[ { "from": "human",
"value": "Saranlah slogan untuk kampanye daur ulang\n" },
{ "from": "gpt",
"value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \
"Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \
"Daur… See the full description on the dataset page: https://huggingface.co/datasets/akahana/alpaca-gpt4-indonesian.indian-government-schemes-2025
Indian Government Schemes Dataset 2026
Dataset Description
The most comprehensive structured dataset of Indian central and state government schemes — 4,693 schemes across all ministries and states, with machine-readable eligibility fields.
Maintained by SmartDuke Technologies · Coimbatore, Tamil Nadu, India
This dataset powers SchemeFit — India's government scheme finder for citizens and businesses.
What Makes This Different
Most existing Indian… See the full description on the dataset page: https://huggingface.co/datasets/har123ish/indian-government-schemes-2025.india2
