CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Ichsan2895 /alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian We wrangled the original dataset format to 'input' & 'output' format. For example: BEFORE: [ { "from": "human", "value": "Saranlah slogan untuk kampanye daur ulang\n" }, { "from": "gpt", "value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \ "Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \ "Daur… See the full description on the dataset page: https://huggingface.co/datasets/Ichsan2895/alpaca-gpt4-indonesian.textquestion-answering10K<n<100K23 likes874 downloads3y agoHugging Face02haipradana /indonesian-twitter-hate-speech-cleaned Dataset Card for indonesian-twitter-hate-speech-cleaned Dataset Summary Cleaned Indonesian Twitter Hate Speech is a curated dataset consisting of Indonesian-language tweets labeled as either hate or neutral. The dataset was collected through a combination of direct scraping from Twitter and aggregation from multiple publicly available GitHub repositories. The data has been cleaned to remove duplicates, irrelevant content, and non-textual noise, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/haipradana/indonesian-twitter-hate-speech-cleaned.texttext-classification10K<n<100K0 likes126 downloads1y agoHugging Face03zaq111 /indonesian-semantic-benchtexttext-classification1K<n<10K0 likes115 downloads1y agoHugging Face04KagChi /indonesian-gambling-words Indonesian Gambling Words Comments Dataset Collector This project provides a Python script for collecting Indonesian YouTube comments. The primary goal is to build a dataset focused on identifying and saving comments that promote online gambling sites. Comments are fetched from specified videos and stored in structured CSV files, organized by a user-defined 'target label'. Features Fetches all comments and replies from a given YouTube video ID. Saves comments to a CSV… See the full description on the dataset page: https://huggingface.co/datasets/KagChi/indonesian-gambling-words.text100K<n<1M0 likes87 downloads1y agoHugging Face05Aarongho /Indonesian-Health-Newstabular10K<n<100K0 likes58 downloads18d agoHugging Face06Faishal-Anwar /alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian We wrangled the original dataset format to 'input' & 'output' format. For example: BEFORE: [ { "from": "human", "value": "Saranlah slogan untuk kampanye daur ulang\n" }, { "from": "gpt", "value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \ "Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \ "Daur… See the full description on the dataset page: https://huggingface.co/datasets/Faishal-Anwar/alpaca-gpt4-indonesian.textquestion-answering10K<n<100K0 likes52 downloads19d agoHugging Face07sepidmnorozy /Indonesian_sentimenttext10K<n<100K2 likes48 downloads4y agoHugging Face08DGurgurov /indonesian_sa Sentiment Analysis Data for the Indonesian Language Dataset Description: This dataset contains a sentiment analysis data from Purwarianti et al. (2019). Data Structure: The data was used for the project on injecting external commonsense knowledge into multilingual Large Language Models. Citation: @inproceedings{purwarianti2019improving, title={Improving bi-lstm performance for indonesian sentiment analysis using paragraph vector}, author={Purwarianti, Ayu and Crisdayanti, Ida… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/indonesian_sa.texttext-classification10K<n<100K1 likes48 downloads2y agoHugging Face09Ichsan2895 /OASST_Top1_IndonesianBase dataset : OpenAssistant/oasst1 We selected the dataset with "english" language & has rank = 1st. Finally, we translate it to Indonesian with Marian NMT and pretrained model from Helsinki-NLP/opus-mt-en-id. CITATION @InProceedings{mariannmt, title = {Marian: Fast Neural Machine Translation in {C++}}, author = {Junczys-Dowmunt, Marcin and Grundkiewicz, Roman and Dwojak, Tomasz and Hoang, Hieu and Heafield, Kenneth and Neckermann… See the full description on the dataset page: https://huggingface.co/datasets/Ichsan2895/OASST_Top1_Indonesian.textquestion-answering1K<n<10K6 likes40 downloads3y agoHugging Face10Speech-data /Indonesian-Speech-Dataset 🎧 Indonesian Speech Dataset The Indonesian Speech Dataset is a high-quality speech audio dataset designed to deliver structured and scalable audio data for AI-powered voice systems. It contains 162 hours of audio data across 821 files, provided in MP3 and WAV formats, with a total size of 210 MB. This well-curated audio dataset ensures balanced and representative voice data, with 51% female and 49% male speakers, and a broad age distribution from 18 to 50+ years. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Indonesian-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes40 downloads6mo agoHugging Face11glhpradipta /alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian We wrangled the original dataset format to 'input' & 'output' format. For example: BEFORE: [ { "from": "human", "value": "Saranlah slogan untuk kampanye daur ulang\n" }, { "from": "gpt", "value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \ "Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \ "Daur… See the full description on the dataset page: https://huggingface.co/datasets/glhpradipta/alpaca-gpt4-indonesian.textquestion-answering10K<n<100K0 likes40 downloads14d agoHugging Face12fadhilakbar /indonesian-names-datasettext10K<n<100K0 likes38 downloads4mo agoHugging Face13egdrga /indonesian-twitter-hate-speech-cleaned Dataset Card for indonesian-twitter-hate-speech-cleaned Dataset Summary Cleaned Indonesian Twitter Hate Speech is a curated dataset consisting of Indonesian-language tweets labeled as either hate or neutral. The dataset was collected through a combination of direct scraping from Twitter and aggregation from multiple publicly available GitHub repositories. The data has been cleaned to remove duplicates, irrelevant content, and non-textual noise, making it suitable… See the full description on the dataset page: https://huggingface.co/datasets/egdrga/indonesian-twitter-hate-speech-cleaned.texttext-classification10K<n<100K0 likes38 downloads7d agoHugging Face14casecrit /2024-indonesian-electionThe dataset encompasses news articles spanning from November 29, 2023, to February 6, 2024, capturing the discourse surrounding the five presidential debates orchestrated by the General Elections Commission. Sourced from reputable platforms such as detik, kompas, and liputan6, the dataset offers a comprehensive insight into the electoral landscape and the media coverage thereof. text10K<n<100K4 likes35 downloads3y agoHugging Face15talithaolga /Indonesian-Emotion-Classificationtexttext-classification10K<n<100K4 likes31 downloads1y agoHugging Face16RimuruChan67 /indonesian-googleplay-sentiment Indonesian Google Play Review Sentiment Dataset A sentiment classification dataset of 5,336 Indonesian-language Google Play reviews collected from three widely used Indonesian applications: Gojek (on-demand transport and delivery), Tokopedia, and Shopee (e-commerce). The dataset accompanies the paper Sentiment Classification of Indonesian Google Play Reviews: A Hybrid Evaluation of Random Forest, CNN, and IndoBERT (ICORIS 2026). Collection Reviews were retrieved… See the full description on the dataset page: https://huggingface.co/datasets/RimuruChan67/indonesian-googleplay-sentiment.texttext-classification1K<n<10K0 likes29 downloads1mo agoHugging Face17pauwdanny /indonesian_hoax_news_oritextn<1K2 likes28 downloads4y agoHugging Face18manueltonneau /indonesian-hate-speech-supersetgated Indonesian Hate Speech Superset This dataset is a superset (N=14,306) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Indonesian hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that: are documented are publicly available focus on hate speech, defined broadly as "any kind of communication in speech, writing or… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/indonesian-hate-speech-superset.tabulartext-classification10K<n<100K4 likes27 downloads2y agoHugging Face19Chemin-AI /Indonesian_DatasetThis dataset is a collection of 50 questions that consists of 4 categories: language, domain, geographical, and combined.Each question has two variations: English & Indonesian. Statistics: Language-Based: 15 Domain-Based: 15 Geographical-Based: 15 Combined: 5 texttranslationn<1K2 likes27 downloads2y agoHugging Face20pauwdanny /indonesian_hoax_news_datasettextn<1K1 likes24 downloads4y agoHugging Face21DGurgurov /indonesian_conceptnet ConceptNet Data for the Indonesian Language Dataset Description: This dataset contains data extracted from ConceptNet using the dedicated module for fetching knowledge from the graph, available on GitHub. Data Structure: The data is converted from triplets into natural text using a pre-defined relationship mapping and split into training and validation sets. It was used for training language adapters for the project aimed at injecting external commonsense knowledge into multilingual… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/indonesian_conceptnet.text10K<n<100K0 likes23 downloads2y agoHugging Face22Hadisawara /indonesian-regional-tax-revenue Indonesian Regional Tax Revenue Analytics Dataset Dataset Description Dataset ini berisi data sintetis realisasi penerimaan pajak daerah di Indonesia, mencakup 10 provinsi, 100+ kota/kabupaten, 13 jenis pajak daerah, dan rentang waktu 7 tahun (2018–2024). Dataset dirancang untuk mendukung penelitian analitik, forecasting, dan klasifikasi di bidang kebijakan fiskal dan tata kelola keuangan daerah. Struktur data terinspirasi dari sistem SIMPAD (Sistem Informasi… See the full description on the dataset page: https://huggingface.co/datasets/Hadisawara/indonesian-regional-tax-revenue.tabulartabular-classification10K<n<100K0 likes22 downloads3mo agoHugging Face23will702 /mendeley-indonesian-stock-sentiment Indonesian Stock Market Sentiment Analysis Dataset Dataset Description This dataset contains Indonesian-language tweets related to stock market sentiment, along with English translations and engagement metrics. It is intended for sentiment analysis tasks on Indonesian financial social media content. Dataset Details Source: Mendeley Data – Indonesian Stock Market Sentiment Analysis HuggingFace: will702/mendeley-indonesian-stock-sentiment File: IDSMSA.csv Rows:… See the full description on the dataset page: https://huggingface.co/datasets/will702/mendeley-indonesian-stock-sentiment.tabulartext-classification1K<n<10K1 likes21 downloads6mo agoHugging Face24ud-synthetic /indonesian-passports Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes. Introduction - Indonesia The Synthetic Indonesia Passports Dataset compiles more than 1,000 AI-generated passport images created for training OCR and computer vision models on identity documents. Each record is fully synthetic, so… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/indonesian-passports.textimage-to-textn<1K3 likes19 downloads2mo agoHugging Face25paranroman /indonesian-deaf-style-emotion-dataset Indonesian Deaf-Style Emotion Dataset Deskripsi Dataset Dataset ini berisi teks Bahasa Indonesia dengan gaya bahasa teman Tuli (Deaf-style) yang dikategorikan ke dalam 6 emosi dasar (Ekman) + 1 label Netral. Dataset ini dirancang untuk melatih model klasifikasi emosi seperti IndoBERT agar lebih inklusif terhadap variasi linguistik komunitas Tuli. Struktur Data Data terdiri dari format CSV dengan kolom utama: text: Kalimat dalam Bahasa Indonesia (gaya bahasa… See the full description on the dataset page: https://huggingface.co/datasets/paranroman/indonesian-deaf-style-emotion-dataset.texttext-classification10K<n<100K0 likes18 downloads8mo agoHugging Face26afrizalha /KamusOne-28M-Indonesian KamusOne (Kamus-1) is a synthethic Indonesian language dataset, generated by Mixtral8x7B. About This dataset was generated by Mixtral 8x7B. For the procedure, Mixtral is instructed that it will act as an Indonesian language dictionary, a native Indonesian speaker, etc. and that it will explain the meaning of a series of Indonesian words. Hence, the name of the dataset ("Kamus", literally "dictionary"). Construction of the word list goes like this. First, we extracted word frequency… See the full description on the dataset page: https://huggingface.co/datasets/afrizalha/KamusOne-28M-Indonesian.texttext-generation100K<n<1M3 likes16 downloads2y agoHugging Face27irfanananda28 /IndonesianEtnoscienceDataset IndonesianEtnoscienceDataset tags: cultural patterns, anthropology, qualitative analysis Note: This is an AI-generated dataset so its content may be inaccurate or false Dataset Description: The 'IndonesianEtnoscienceDataset' is a collection of qualitative data gathered from various sources including interviews, ethnographic studies, and anthropological research on the indigenous knowledge systems in Indonesia. The dataset focuses on the cultural patterns and practices that are… See the full description on the dataset page: https://huggingface.co/datasets/irfanananda28/IndonesianEtnoscienceDataset.textn<1K0 likes16 downloads1y agoHugging Face28jojo-ai-mst /Roleplay-Indonesian RolePlay-Indonesian Roleplay-Indonesian Dataset is a dataset for roleplaying in the Indonesian language for Large Language Model. The base dataset is GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API. For more information and other language datasets for roleplay, it can be found at this… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Indonesian.texttext-generation1K<n<10K2 likes14 downloads2y agoHugging Face29ermandmand /indonesian-app-feedback-datasettextn<1K0 likes14 downloads8mo agoHugging Face30Hemg /indonesian2englishtext100K<n<1M0 likes13 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.