datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian
We wrangled the original dataset format to 'input' & 'output' format. For example:
BEFORE:
[ { "from": "human",
"value": "Saranlah slogan untuk kampanye daur ulang\n" },
{ "from": "gpt",
"value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \
"Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \
"Daur… See the full description on the dataset page: https://huggingface.co/datasets/Ichsan2895/alpaca-gpt4-indonesian.Indonesia-Stock-Symbols-and-Metadata
Indonesia Stock Symbols & Company Metadata
This dataset contains stock symbols and basic company metadata for all listed companies in Indonesia.It is updated weekly if new changes are there.
📊 Dataset Contents
The dataset is provided as a CSV file with the following columns:
Column
Description
name
Full company name
ticker
Stock ticker symbol (e.g., AAPL, MSFT)
market
The exchange/market where the stock is listed
sector
The primary business sector of… See the full description on the dataset page: https://huggingface.co/datasets/kjhq/Indonesia-Stock-Symbols-and-Metadata.indonesia-slangtwitter_indonesia_sarcastic
Twitter Indonesia Sarcastic
Twitter Indonesia Sarcastic is a dataset intended for sarcasm detection in the Indonesian language. This dataset is introduced in Khotijah et al. (2020), whereby Indonesian tweets are collected and labeled as either sarcastic or non-sarcastic. We took the raw data, and performed several cleaning procedures such as: sentence order re-reversal, deduplication with minHash LSH, PII masking to remove usernames, hashtags, emails, URLs, and finally a random… See the full description on the dataset page: https://huggingface.co/datasets/w11wo/twitter_indonesia_sarcastic.indonesian-twitter-hate-speech-cleaned
Dataset Card for indonesian-twitter-hate-speech-cleaned
Dataset Summary
Cleaned Indonesian Twitter Hate Speech is a curated dataset consisting of Indonesian-language tweets labeled as either hate or neutral. The dataset was collected through a combination of direct scraping from Twitter and aggregation from multiple publicly available GitHub repositories.
The data has been cleaned to remove duplicates, irrelevant content, and non-textual noise, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/haipradana/indonesian-twitter-hate-speech-cleaned.indonesian-semantic-benchindonesian-gambling-words
Indonesian Gambling Words Comments Dataset Collector
This project provides a Python script for collecting Indonesian YouTube comments. The primary goal is to build a dataset focused on identifying and saving comments that promote online gambling sites. Comments are fetched from specified videos and stored in structured CSV files, organized by a user-defined 'target label'.
Features
Fetches all comments and replies from a given YouTube video ID.
Saves comments to a CSV… See the full description on the dataset page: https://huggingface.co/datasets/KagChi/indonesian-gambling-words.stif-indonesia
Dataset Card for "stif-indonesia"
STIF-Indonesia
A dataset of "Semi-Supervised Low-Resource Style Transfer of Indonesian Informal to Formal Language with Iterative Forward-Translation".
You can also find Indonesian informal-formal parallel corpus in this repository.
Description
We were researching transforming a sentence from informal to its formal form. Our work addresses a style-transfer from informal to formal Indonesian as a low-resource machine… See the full description on the dataset page: https://huggingface.co/datasets/haryoaw/stif-indonesia.kamus-besar-bahasa-indonesiaIndonesian-Health-Newsalpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian
We wrangled the original dataset format to 'input' & 'output' format. For example:
BEFORE:
[ { "from": "human",
"value": "Saranlah slogan untuk kampanye daur ulang\n" },
{ "from": "gpt",
"value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \
"Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \
"Daur… See the full description on the dataset page: https://huggingface.co/datasets/Faishal-Anwar/alpaca-gpt4-indonesian.Indonesian_sentimentIndonesia-Stock-Symbols-and-Metadata
Indonesia Stock Symbols & Company Metadata
This dataset contains stock symbols and basic company metadata for all listed companies in Indonesia.It is updated weekly if new changes are there.
📊 Dataset Contents
The dataset is provided as a CSV file with the following columns:
Column
Description
name
Full company name
ticker
Stock ticker symbol (e.g., AAPL, MSFT)
market
The exchange/market where the stock is listed
sector
The primary business sector of… See the full description on the dataset page: https://huggingface.co/datasets/DimmsMath/Indonesia-Stock-Symbols-and-Metadata.indonesian_sa
Sentiment Analysis Data for the Indonesian Language
Dataset Description:
This dataset contains a sentiment analysis data from Purwarianti et al. (2019).
Data Structure:
The data was used for the project on injecting external commonsense knowledge into multilingual Large Language Models.
Citation:
@inproceedings{purwarianti2019improving,
title={Improving bi-lstm performance for indonesian sentiment analysis using paragraph vector},
author={Purwarianti, Ayu and Crisdayanti, Ida… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/indonesian_sa.OASST_Top1_IndonesianBase dataset : OpenAssistant/oasst1
We selected the dataset with "english" language & has rank = 1st.
Finally, we translate it to Indonesian with Marian NMT and pretrained model from Helsinki-NLP/opus-mt-en-id.
CITATION
@InProceedings{mariannmt,
title = {Marian: Fast Neural Machine Translation in {C++}},
author = {Junczys-Dowmunt, Marcin and Grundkiewicz, Roman and
Dwojak, Tomasz and Hoang, Hieu and Heafield, Kenneth and
Neckermann… See the full description on the dataset page: https://huggingface.co/datasets/Ichsan2895/OASST_Top1_Indonesian.alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian
We wrangled the original dataset format to 'input' & 'output' format. For example:
BEFORE:
[ { "from": "human",
"value": "Saranlah slogan untuk kampanye daur ulang\n" },
{ "from": "gpt",
"value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \
"Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \
"Daur… See the full description on the dataset page: https://huggingface.co/datasets/glhpradipta/alpaca-gpt4-indonesian.indonesian-names-datasetindonesian-twitter-hate-speech-cleaned
Dataset Card for indonesian-twitter-hate-speech-cleaned
Dataset Summary
Cleaned Indonesian Twitter Hate Speech is a curated dataset consisting of Indonesian-language tweets labeled as either hate or neutral. The dataset was collected through a combination of direct scraping from Twitter and aggregation from multiple publicly available GitHub repositories.
The data has been cleaned to remove duplicates, irrelevant content, and non-textual noise, making it suitable… See the full description on the dataset page: https://huggingface.co/datasets/egdrga/indonesian-twitter-hate-speech-cleaned.Indonesian-Speech-Dataset
🎧 Indonesian Speech Dataset
The Indonesian Speech Dataset is a high-quality speech audio dataset designed to deliver structured and scalable audio data for AI-powered voice systems. It contains 162 hours of audio data across 821 files, provided in MP3 and WAV formats, with a total size of 210 MB. This well-curated audio dataset ensures balanced and representative voice data, with 51% female and 49% male speakers, and a broad age distribution from 18 to 50+ years. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Indonesian-Speech-Dataset.indonesia-medical-qna
indonesia-medical-qna
This dataset is configured so the Hugging Face Dataset Viewer loads qna.csv as the primary data file.
The repository also contains all_links.csv, which has a different schema and is kept as an auxiliary file rather than part of the default viewer configuration.
This avoids schema-casting errors caused by the Hub trying to combine both CSV files into one split.
2024-indonesian-electionThe dataset encompasses news articles spanning from November 29, 2023, to February 6, 2024, capturing the discourse surrounding the five presidential debates orchestrated by the General Elections Commission. Sourced from reputable platforms such as detik, kompas, and liputan6, the dataset offers a comprehensive insight into the electoral landscape and the media coverage thereof.
indonesian-hate-speech-superset
Indonesian Hate Speech Superset
This dataset is a superset (N=14,306) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Indonesian hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available
focus on hate speech, defined broadly as "any kind of communication in speech, writing or… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/indonesian-hate-speech-superset.Indonesian-Emotion-Classificationindonesian-googleplay-sentiment
Indonesian Google Play Review Sentiment Dataset
A sentiment classification dataset of 5,336 Indonesian-language Google Play reviews
collected from three widely used Indonesian applications: Gojek (on-demand transport
and delivery), Tokopedia, and Shopee (e-commerce).
The dataset accompanies the paper Sentiment Classification of Indonesian Google Play
Reviews: A Hybrid Evaluation of Random Forest, CNN, and IndoBERT (ICORIS 2026).
Collection
Reviews were retrieved… See the full description on the dataset page: https://huggingface.co/datasets/RimuruChan67/indonesian-googleplay-sentiment.indonesian_hoax_news_oriIndonesian_DatasetThis dataset is a collection of 50 questions that consists of 4 categories: language, domain, geographical, and combined.Each question has two variations: English & Indonesian.
Statistics:
Language-Based: 15
Domain-Based: 15
Geographical-Based: 15
Combined: 5
indonesian_conceptnet
ConceptNet Data for the Indonesian Language
Dataset Description:
This dataset contains data extracted from ConceptNet using the dedicated module for fetching knowledge from the graph, available on GitHub.
Data Structure:
The data is converted from triplets into natural text using a pre-defined relationship mapping and split into training and validation sets. It was used for training language adapters for the project aimed at injecting external commonsense knowledge into multilingual… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/indonesian_conceptnet.indonesia-affordable-housing
Indonesia Affordable Housing Dataset
A comprehensive dataset of affordable housing projects across Indonesia, containing detailed information about residential properties, specifications, locations, pricing, and developer data.
Dataset Creator: web3hungry
Dataset ID: web3hungry/indonesia-affordable-housing
License: CC0 1.0
Dataset Overview
This dataset provides extensive data on housing developments throughout Indonesia, covering both subsidized and commercial… See the full description on the dataset page: https://huggingface.co/datasets/web3hungry/indonesia-affordable-housing.indonesian-regional-tax-revenue
Indonesian Regional Tax Revenue Analytics Dataset
Dataset Description
Dataset ini berisi data sintetis realisasi penerimaan pajak daerah di Indonesia, mencakup 10 provinsi, 100+ kota/kabupaten, 13 jenis pajak daerah, dan rentang waktu 7 tahun (2018–2024). Dataset dirancang untuk mendukung penelitian analitik, forecasting, dan klasifikasi di bidang kebijakan fiskal dan tata kelola keuangan daerah.
Struktur data terinspirasi dari sistem SIMPAD (Sistem Informasi… See the full description on the dataset page: https://huggingface.co/datasets/Hadisawara/indonesian-regional-tax-revenue.indonesian_hoax_news_dataset
