datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
t5gemma2-indonesia-instruct-v1
T5Gemma-2 Indonesian Instruct — Mono-Repo
Satu repositori dataset HF untuk seluruh data pelatihan T5-Gemma-2 bahasa Indonesia.
Diorganisasi per fungsi (fondasi → spesifik → preferensi) dengan folder/subfolder,
setiap config = folder dan berisi split train + validation (80:20) di level percakapan.
Struktur (by fungsi)
t5gemma2-indonesia-instruct-v1/
├── README.md
├── manifest.json
├── chat_idx_map.json
├── foundation/ ← FASE 1 · fondasi Bahasa… See the full description on the dataset page: https://huggingface.co/datasets/daruokta/t5gemma2-indonesia-instruct-v1.Indonesian-ASR-11-Class-Dataset
Indonesian ASR 11-Class Dataset
Public Hugging Face repository for an Indonesian ASR corpus and its paper-supporting benchmark artifacts.
Dataset summary
Audio files: 104,500 WAV files
Real/human recordings: 104,368
Synthetic repair files: 132
Sentence classes: 11 Indonesian sentence categories
Canonical balanced sentence slots: 209 (11 categories × 19 retained slots)
Public speaker labels: M1..M12, F1..F8, plus synthetic labels Ms*/Fs*
Audio format: 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/Atika88/Indonesian-ASR-11-Class-Dataset.reddit_indonesia_sarcastic
Reddit Indonesia Sarcastic
Reddit Indonesia Sarcastic is a dataset intended for sarcasm detection in the Indonesian language. This dataset is inspired by the data collection procedure introduced in Ranti, K.S., & Girsang, A.S (2020), whereby Reddit comments from r/indonesia subreddit are collected and filtered by the existence of an /s tag at the end of the comment. We collected Reddit comments from 2020-01 to 2023-09 from Academic Torrents and applied the aforementioned procedure.… See the full description on the dataset page: https://huggingface.co/datasets/w11wo/reddit_indonesia_sarcastic.laion-2b-indonesia-subsetindonesian-recipes
Resep Masakan Indonesia 🍛
Kumpulan resep masakan Indonesia autentik — dari rendang sampai es cendol, lengkap dengan bahan, langkah, tingkat kesulitan, waktu, dan daerah asal.
Kenapa dataset ini ada?
Resep adalah salah satu konten paling dicari untuk LLM (assistant masak) — tapi dataset resep Indonesia di HF nyaris kosong (cuma 1 yang 34 likes). Gw isi gap itu dengan resep-resep yang benar-benar asli Indonesia, bukan versi western yang diterjemahkan.… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-recipes.ipfs_indonesia_laws_ir
Indonesia legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_indonesia_laws (revision 796ca261e322f0325e210285660fa944654dea07) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Indonesia prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_indonesia_laws_ir.t5gemma2-indonesia-chat-formatted
T5Gemma-2 Indonesian Chat & QA Dataset
A high-quality Indonesian language multi-turn conversation and reading comprehension dataset, specifically formatted for instruction tuning of sequence-to-sequence (Seq2Seq) models like T5-Gemma / T5-Gemma-2.
Dataset Description
This dataset contains over 7,400 multi-turn conversations and document-based Q&A in Bahasa Indonesia. It covers diverse topics including everyday life, technology, general knowledge, and structured… See the full description on the dataset page: https://huggingface.co/datasets/daruokta/t5gemma2-indonesia-chat-formatted.Indonesian-Health-Newsfineweb_edu_indonesia-100kIndonesian-High-School-Basketball-Complete-StatsDemographics told you who they are. This tells you how they play.
A structured, multi-table dataset covering high school basketball competition in Indonesia match results with full team box stats, per-player per-game performance, and multi-season player statistics. Delivered as compressed Parquet files, ready for pandas, polars, DuckDB, or the datasets library.
This is the follow-up to illustricious/indonesian-highschool-basketball-demographics(player height, weight, gender, and school city… See the full description on the dataset page: https://huggingface.co/datasets/illustricious/Indonesian-High-School-Basketball-Complete-Stats.apbd-pemda-indonesiaindonesian-emotion
Emosi Bahasa Indonesia 💢
Dataset klasifikasi emosi berbahasa Indonesia — 267 teks natural dengan label 7 emosi + intensitas + konteks.
Kenapa dataset ini ada?
Sentiment analysis Bahasa Indonesia ada, tapi emotion classification (7 emosi) belum ada di HF. Sentiment cuma bilang positif/negatif/netral — emosi lebih dalam: marah vs jijik vs takut itu beda, dan buat chatbot/analisis medsos, label emosi jauh lebih berguna.
Isi
Field
Tipe
Contoh… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-emotion.indonesian-hate-speech-superset
Indonesian Hate Speech Superset
This dataset is a superset (N=14,306) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Indonesian hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available
focus on hate speech, defined broadly as "any kind of communication in speech, writing or… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/indonesian-hate-speech-superset.Indonesian-STEM-Textbook-DatasetDataset Description:
This dataset is a large-scale collection of Indonesian STEM textbook data, containing 5,169 books and 208.30 million words, designed to support the development and training of advanced NLP systems and AI models for scientific understanding, problem-solving, and concept learning in Bahasa.
Full Dataset Overview
This dataset is part of a large-scale multilingual educational corpus containing over 3+ billion words across 5,000+ subjects, supported by interwoven images for… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Indonesian-STEM-Textbook-Dataset.XSum-Indonesia-with-Entailment-Labelindonesian-highschool-basketball-demographics
Indonesia Youth Basketball Players: Physical & Demographic Dataset
Dataset Summary
This dataset contains cleaned, anonymized physical and demographic profiles of youth basketball players across various provinces in Indonesia. It provides valuable insights into the anthropometric characteristics (height, weight, age) and geographic distribution of young Indonesian athletes.
All Personally Identifiable Information (PII)—including player names, specific city/school… See the full description on the dataset page: https://huggingface.co/datasets/illustricious/indonesian-highschool-basketball-demographics.XSUM-Indonesia-AMR-NLI
XSUM-Indonesia-AMR-NLI
Deskripsi
Dataset ini berisi kumpulan data dalam bahasa Indonesia yang dirancang untuk tugas Natural Language Inference (NLI).
Setiap instans data terdiri dari
teks sumber (source_text) yang ada pada dataset XSum
teks yang dihasilkan (generated_indonesian yang berfungsi sebagai hipotesis dibuat dari AMR Perturbasi)
skor (score) yang menunjukkan hubungan antara keduanya (0 untuk non-entailment, 1 untuk entailment).
ringkasan asli (target_summary)… See the full description on the dataset page: https://huggingface.co/datasets/fabhiansan/XSUM-Indonesia-AMR-NLI.XSum-Indonesia-with-PerturbationIndonesian-Non-STEM-Textbook-DatasetDataset Description:
This dataset is a large-scale collection of Indonesian Non-STEM textbook data, containing 4,098 books and 182.10 million words, designed to support the development and training of advanced NLP systems and AI models for language understanding, reasoning, and general knowledge learning in Bahasa.
Full Dataset Overview
This dataset is part of a large-scale multilingual educational corpus containing over 3+ billion words across 5,000+ subjects, supported by interwoven images… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Indonesian-Non-STEM-Textbook-Dataset.indonesia-affordable-housing
Indonesia Affordable Housing Dataset
A comprehensive dataset of affordable housing projects across Indonesia, containing detailed information about residential properties, specifications, locations, pricing, and developer data.
Dataset Creator: web3hungry
Dataset ID: web3hungry/indonesia-affordable-housing
License: CC0 1.0
Dataset Overview
This dataset provides extensive data on housing developments throughout Indonesia, covering both subsidized and commercial… See the full description on the dataset page: https://huggingface.co/datasets/web3hungry/indonesia-affordable-housing.indonesian-regional-tax-revenue
Indonesian Regional Tax Revenue Analytics Dataset
Dataset Description
Dataset ini berisi data sintetis realisasi penerimaan pajak daerah di Indonesia, mencakup 10 provinsi, 100+ kota/kabupaten, 13 jenis pajak daerah, dan rentang waktu 7 tahun (2018–2024). Dataset dirancang untuk mendukung penelitian analitik, forecasting, dan klasifikasi di bidang kebijakan fiskal dan tata kelola keuangan daerah.
Struktur data terinspirasi dari sistem SIMPAD (Sistem Informasi… See the full description on the dataset page: https://huggingface.co/datasets/Hadisawara/indonesian-regional-tax-revenue.asia-demographics-dhs-data-for-indonesia
Indonesia - National Demographic and Health Data
Contains data from the DHS data portal. There is also a dataset containing Indonesia - Subnational Demographic and Health Data on HDX. The DHS Program Application Programming Interface (API) provid
Resources
Resource Count: 40
Formats: csv
Last Updated: 2026-04-20
Coverage
idn
Tags
demographics, health
License
License ID: hdx-other
Citation
@dataset{{{slug}},
title =… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-demographics-dhs-data-for-indonesia.indonesian-honorifics
Panggilan Hormat Nusantara: Indonesian Honorifics
Gelar dan kata panggilan hormat yang benar-benar dipakai orang Indonesia — dari "Pak" dan "Bu" yang ada di mana-mana, sampai "Ompung" di Tanah Batak, "Daeng" di Makassar, "Raden" di Jawa, dan "Tengku" di Melayu. Setiap panggilan membawa cerita: ada yang soal keluarga, ada yang soal bangsawan, ada yang soal agama, ada yang soal dagang di pasar.
Isi dataset
120 panggilan hormat dengan kolom:
kolom
arti
id… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-honorifics.asia-aid-flows-iati-indonesia
Indonesia - Current IATI Aid Activities
Publisher: International Aid Transparency Initiative · Source: HDX · License: hdx-other · Updated: 2026-05-06
Abstract
List of active aid activities shared via the International Aid Transparency Initiative (IATI). Includes both humanitarian and development activities. More information on each activity (including financial data) is available from http://www.d-portal.org
Each row in this dataset represents country-level aggregates.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-aid-flows-iati-indonesia.indonesian-product-reviews
Ulasan Produk Indonesia 🛒
Kumpulan 158 ulasan produk bahasa Indonesia — gaya orang biasa nulis review beneran: santai, detail, kadang emoji. 7 kategori produk, rating 1-5, sentiment berlabel.
Kenapa dataset ini ada?
Sentiment analysis e-commerce Indonesia butuh data ulasan yang realistis — bukan teks formal. Dataset ulasan produk Indonesia di HF belum ada (turkish_product_reviews 291 dl — bukti permintaan). Ini yang pertama: ulasan natural dengan slang & grammar… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-product-reviews.OASST1-INDONESIANglobalopinionv3_indonesiaindonesian-criminal-law-statutorymendeley-indonesian-stock-sentiment
Indonesian Stock Market Sentiment Analysis Dataset
Dataset Description
This dataset contains Indonesian-language tweets related to stock market sentiment, along with English translations and engagement metrics. It is intended for sentiment analysis tasks on Indonesian financial social media content.
Dataset Details
Source: Mendeley Data – Indonesian Stock Market Sentiment Analysis
HuggingFace: will702/mendeley-indonesian-stock-sentiment
File: IDSMSA.csv
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/will702/mendeley-indonesian-stock-sentiment.indonesian-recipes
Indonesian Recipes
A structured collection of Indonesian recipes for fine-tuning text-generation models. Each row is a single recipe with a title, an ingredient list, and ordered preparation steps.
Schema
Column
Type
Description
title
string
Recipe name
ingredients
list<string>
One item per ingredient line
steps
list<string>
Ordered preparation steps
num_ingredients
int
len(ingredients)
num_steps
int
len(steps)
char_count
int
Total characters… See the full description on the dataset page: https://huggingface.co/datasets/junwatu/indonesian-recipes.
