datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian
We wrangled the original dataset format to 'input' & 'output' format. For example:
BEFORE:
[ { "from": "human",
"value": "Saranlah slogan untuk kampanye daur ulang\n" },
{ "from": "gpt",
"value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \
"Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \
"Daur… See the full description on the dataset page: https://huggingface.co/datasets/Ichsan2895/alpaca-gpt4-indonesian.indonesian-twitter-hate-speech-cleaned
Dataset Card for indonesian-twitter-hate-speech-cleaned
Dataset Summary
Cleaned Indonesian Twitter Hate Speech is a curated dataset consisting of Indonesian-language tweets labeled as either hate or neutral. The dataset was collected through a combination of direct scraping from Twitter and aggregation from multiple publicly available GitHub repositories.
The data has been cleaned to remove duplicates, irrelevant content, and non-textual noise, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/haipradana/indonesian-twitter-hate-speech-cleaned.indonesian-semantic-benchindonesian-gambling-words
Indonesian Gambling Words Comments Dataset Collector
This project provides a Python script for collecting Indonesian YouTube comments. The primary goal is to build a dataset focused on identifying and saving comments that promote online gambling sites. Comments are fetched from specified videos and stored in structured CSV files, organized by a user-defined 'target label'.
Features
Fetches all comments and replies from a given YouTube video ID.
Saves comments to a CSV… See the full description on the dataset page: https://huggingface.co/datasets/KagChi/indonesian-gambling-words.Indonesian-Health-Newsalpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian
We wrangled the original dataset format to 'input' & 'output' format. For example:
BEFORE:
[ { "from": "human",
"value": "Saranlah slogan untuk kampanye daur ulang\n" },
{ "from": "gpt",
"value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \
"Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \
"Daur… See the full description on the dataset page: https://huggingface.co/datasets/Faishal-Anwar/alpaca-gpt4-indonesian.Indonesian_sentimentindonesian_sa
Sentiment Analysis Data for the Indonesian Language
Dataset Description:
This dataset contains a sentiment analysis data from Purwarianti et al. (2019).
Data Structure:
The data was used for the project on injecting external commonsense knowledge into multilingual Large Language Models.
Citation:
@inproceedings{purwarianti2019improving,
title={Improving bi-lstm performance for indonesian sentiment analysis using paragraph vector},
author={Purwarianti, Ayu and Crisdayanti, Ida… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/indonesian_sa.OASST_Top1_IndonesianBase dataset : OpenAssistant/oasst1
We selected the dataset with "english" language & has rank = 1st.
Finally, we translate it to Indonesian with Marian NMT and pretrained model from Helsinki-NLP/opus-mt-en-id.
CITATION
@InProceedings{mariannmt,
title = {Marian: Fast Neural Machine Translation in {C++}},
author = {Junczys-Dowmunt, Marcin and Grundkiewicz, Roman and
Dwojak, Tomasz and Hoang, Hieu and Heafield, Kenneth and
Neckermann… See the full description on the dataset page: https://huggingface.co/datasets/Ichsan2895/OASST_Top1_Indonesian.Indonesian-Speech-Dataset
🎧 Indonesian Speech Dataset
The Indonesian Speech Dataset is a high-quality speech audio dataset designed to deliver structured and scalable audio data for AI-powered voice systems. It contains 162 hours of audio data across 821 files, provided in MP3 and WAV formats, with a total size of 210 MB. This well-curated audio dataset ensures balanced and representative voice data, with 51% female and 49% male speakers, and a broad age distribution from 18 to 50+ years. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Indonesian-Speech-Dataset.alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian
We wrangled the original dataset format to 'input' & 'output' format. For example:
BEFORE:
[ { "from": "human",
"value": "Saranlah slogan untuk kampanye daur ulang\n" },
{ "from": "gpt",
"value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \
"Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \
"Daur… See the full description on the dataset page: https://huggingface.co/datasets/glhpradipta/alpaca-gpt4-indonesian.indonesian-names-datasetindonesian-twitter-hate-speech-cleaned
Dataset Card for indonesian-twitter-hate-speech-cleaned
Dataset Summary
Cleaned Indonesian Twitter Hate Speech is a curated dataset consisting of Indonesian-language tweets labeled as either hate or neutral. The dataset was collected through a combination of direct scraping from Twitter and aggregation from multiple publicly available GitHub repositories.
The data has been cleaned to remove duplicates, irrelevant content, and non-textual noise, making it suitable… See the full description on the dataset page: https://huggingface.co/datasets/egdrga/indonesian-twitter-hate-speech-cleaned.2024-indonesian-electionThe dataset encompasses news articles spanning from November 29, 2023, to February 6, 2024, capturing the discourse surrounding the five presidential debates orchestrated by the General Elections Commission. Sourced from reputable platforms such as detik, kompas, and liputan6, the dataset offers a comprehensive insight into the electoral landscape and the media coverage thereof.
Indonesian-Emotion-Classificationindonesian-googleplay-sentiment
Indonesian Google Play Review Sentiment Dataset
A sentiment classification dataset of 5,336 Indonesian-language Google Play reviews
collected from three widely used Indonesian applications: Gojek (on-demand transport
and delivery), Tokopedia, and Shopee (e-commerce).
The dataset accompanies the paper Sentiment Classification of Indonesian Google Play
Reviews: A Hybrid Evaluation of Random Forest, CNN, and IndoBERT (ICORIS 2026).
Collection
Reviews were retrieved… See the full description on the dataset page: https://huggingface.co/datasets/RimuruChan67/indonesian-googleplay-sentiment.indonesian_hoax_news_oriindonesian-hate-speech-superset
Indonesian Hate Speech Superset
This dataset is a superset (N=14,306) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Indonesian hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available
focus on hate speech, defined broadly as "any kind of communication in speech, writing or… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/indonesian-hate-speech-superset.Indonesian_DatasetThis dataset is a collection of 50 questions that consists of 4 categories: language, domain, geographical, and combined.Each question has two variations: English & Indonesian.
Statistics:
Language-Based: 15
Domain-Based: 15
Geographical-Based: 15
Combined: 5
indonesian_hoax_news_datasetindonesian_conceptnet
ConceptNet Data for the Indonesian Language
Dataset Description:
This dataset contains data extracted from ConceptNet using the dedicated module for fetching knowledge from the graph, available on GitHub.
Data Structure:
The data is converted from triplets into natural text using a pre-defined relationship mapping and split into training and validation sets. It was used for training language adapters for the project aimed at injecting external commonsense knowledge into multilingual… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/indonesian_conceptnet.indonesian-regional-tax-revenue
Indonesian Regional Tax Revenue Analytics Dataset
Dataset Description
Dataset ini berisi data sintetis realisasi penerimaan pajak daerah di Indonesia, mencakup 10 provinsi, 100+ kota/kabupaten, 13 jenis pajak daerah, dan rentang waktu 7 tahun (2018–2024). Dataset dirancang untuk mendukung penelitian analitik, forecasting, dan klasifikasi di bidang kebijakan fiskal dan tata kelola keuangan daerah.
Struktur data terinspirasi dari sistem SIMPAD (Sistem Informasi… See the full description on the dataset page: https://huggingface.co/datasets/Hadisawara/indonesian-regional-tax-revenue.mendeley-indonesian-stock-sentiment
Indonesian Stock Market Sentiment Analysis Dataset
Dataset Description
This dataset contains Indonesian-language tweets related to stock market sentiment, along with English translations and engagement metrics. It is intended for sentiment analysis tasks on Indonesian financial social media content.
Dataset Details
Source: Mendeley Data – Indonesian Stock Market Sentiment Analysis
HuggingFace: will702/mendeley-indonesian-stock-sentiment
File: IDSMSA.csv
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/will702/mendeley-indonesian-stock-sentiment.indonesian-passports
Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes.
Introduction - Indonesia
The Synthetic Indonesia Passports Dataset compiles more than 1,000 AI-generated passport images created for training OCR and computer vision models on identity documents. Each record is fully synthetic, so… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/indonesian-passports.indonesian-deaf-style-emotion-dataset
Indonesian Deaf-Style Emotion Dataset
Deskripsi Dataset
Dataset ini berisi teks Bahasa Indonesia dengan gaya bahasa teman Tuli (Deaf-style) yang dikategorikan ke dalam 6 emosi dasar (Ekman) + 1 label Netral. Dataset ini dirancang untuk melatih model klasifikasi emosi seperti IndoBERT agar lebih inklusif terhadap variasi linguistik komunitas Tuli.
Struktur Data
Data terdiri dari format CSV dengan kolom utama:
text: Kalimat dalam Bahasa Indonesia (gaya bahasa… See the full description on the dataset page: https://huggingface.co/datasets/paranroman/indonesian-deaf-style-emotion-dataset.KamusOne-28M-Indonesian
KamusOne (Kamus-1) is a synthethic Indonesian language dataset, generated by Mixtral8x7B.
About
This dataset was generated by Mixtral 8x7B. For the procedure, Mixtral is instructed that it will act as an Indonesian language dictionary, a native Indonesian speaker, etc. and that it will explain the meaning of a series of Indonesian words. Hence, the name of the dataset ("Kamus", literally "dictionary"). Construction of the word list goes like this. First, we extracted word frequency… See the full description on the dataset page: https://huggingface.co/datasets/afrizalha/KamusOne-28M-Indonesian.IndonesianEtnoscienceDataset
IndonesianEtnoscienceDataset
tags: cultural patterns, anthropology, qualitative analysis
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'IndonesianEtnoscienceDataset' is a collection of qualitative data gathered from various sources including interviews, ethnographic studies, and anthropological research on the indigenous knowledge systems in Indonesia. The dataset focuses on the cultural patterns and practices that are… See the full description on the dataset page: https://huggingface.co/datasets/irfanananda28/IndonesianEtnoscienceDataset.Roleplay-Indonesian
RolePlay-Indonesian
Roleplay-Indonesian Dataset is a dataset for roleplaying in the Indonesian language for Large Language Model.
The base dataset is GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API.
For more information and other language datasets for roleplay, it can be found at this… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Indonesian.indonesian-app-feedback-datasetindonesian2english
