CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Ichsan2895 /alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian We wrangled the original dataset format to 'input' & 'output' format. For example: BEFORE: [ { "from": "human", "value": "Saranlah slogan untuk kampanye daur ulang\n" }, { "from": "gpt", "value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \ "Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \ "Daur… See the full description on the dataset page: https://huggingface.co/datasets/Ichsan2895/alpaca-gpt4-indonesian.textquestion-answering10K<n<100K23 likes861 downloads3y agoHugging Face02kjhq /Indonesia-Stock-Symbols-and-Metadata Indonesia Stock Symbols & Company Metadata This dataset contains stock symbols and basic company metadata for all listed companies in Indonesia.It is updated weekly if new changes are there. 📊 Dataset Contents The dataset is provided as a CSV file with the following columns: Column Description name Full company name ticker Stock ticker symbol (e.g., AAPL, MSFT) market The exchange/market where the stock is listed sector The primary business sector of… See the full description on the dataset page: https://huggingface.co/datasets/kjhq/Indonesia-Stock-Symbols-and-Metadata.textn<1K0 likes554 downloads1y agoHugging Face03theonlydo /indonesia-slangtext1K<n<10K13 likes198 downloads3y agoHugging Face04w11wo /twitter_indonesia_sarcastic Twitter Indonesia Sarcastic Twitter Indonesia Sarcastic is a dataset intended for sarcasm detection in the Indonesian language. This dataset is introduced in Khotijah et al. (2020), whereby Indonesian tweets are collected and labeled as either sarcastic or non-sarcastic. We took the raw data, and performed several cleaning procedures such as: sentence order re-reversal, deduplication with minHash LSH, PII masking to remove usernames, hashtags, emails, URLs, and finally a random… See the full description on the dataset page: https://huggingface.co/datasets/w11wo/twitter_indonesia_sarcastic.text1K<n<10K10 likes162 downloads3y agoHugging Face05haipradana /indonesian-twitter-hate-speech-cleaned Dataset Card for indonesian-twitter-hate-speech-cleaned Dataset Summary Cleaned Indonesian Twitter Hate Speech is a curated dataset consisting of Indonesian-language tweets labeled as either hate or neutral. The dataset was collected through a combination of direct scraping from Twitter and aggregation from multiple publicly available GitHub repositories. The data has been cleaned to remove duplicates, irrelevant content, and non-textual noise, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/haipradana/indonesian-twitter-hate-speech-cleaned.texttext-classification10K<n<100K0 likes121 downloads1y agoHugging Face06zaq111 /indonesian-semantic-benchtexttext-classification1K<n<10K0 likes115 downloads1y agoHugging Face07KagChi /indonesian-gambling-words Indonesian Gambling Words Comments Dataset Collector This project provides a Python script for collecting Indonesian YouTube comments. The primary goal is to build a dataset focused on identifying and saving comments that promote online gambling sites. Comments are fetched from specified videos and stored in structured CSV files, organized by a user-defined 'target label'. Features Fetches all comments and replies from a given YouTube video ID. Saves comments to a CSV… See the full description on the dataset page: https://huggingface.co/datasets/KagChi/indonesian-gambling-words.text100K<n<1M0 likes86 downloads1y agoHugging Face08haryoaw /stif-indonesia Dataset Card for "stif-indonesia" STIF-Indonesia A dataset of "Semi-Supervised Low-Resource Style Transfer of Indonesian Informal to Formal Language with Iterative Forward-Translation". You can also find Indonesian informal-formal parallel corpus in this repository. Description We were researching transforming a sentence from informal to its formal form. Our work addresses a style-transfer from informal to formal Indonesian as a low-resource machine… See the full description on the dataset page: https://huggingface.co/datasets/haryoaw/stif-indonesia.texttranslation1K<n<10K11 likes83 downloads3y agoHugging Face09Lyon28 /kamus-besar-bahasa-indonesiatexttext-generation100K<n<1M2 likes75 downloads1y agoHugging Face10Aarongho /Indonesian-Health-Newstabular10K<n<100K0 likes55 downloads19d agoHugging Face11Faishal-Anwar /alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian We wrangled the original dataset format to 'input' & 'output' format. For example: BEFORE: [ { "from": "human", "value": "Saranlah slogan untuk kampanye daur ulang\n" }, { "from": "gpt", "value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \ "Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \ "Daur… See the full description on the dataset page: https://huggingface.co/datasets/Faishal-Anwar/alpaca-gpt4-indonesian.textquestion-answering10K<n<100K0 likes53 downloads20d agoHugging Face12sepidmnorozy /Indonesian_sentimenttext10K<n<100K2 likes48 downloads4y agoHugging Face13DimmsMath /Indonesia-Stock-Symbols-and-Metadata Indonesia Stock Symbols & Company Metadata This dataset contains stock symbols and basic company metadata for all listed companies in Indonesia.It is updated weekly if new changes are there. 📊 Dataset Contents The dataset is provided as a CSV file with the following columns: Column Description name Full company name ticker Stock ticker symbol (e.g., AAPL, MSFT) market The exchange/market where the stock is listed sector The primary business sector of… See the full description on the dataset page: https://huggingface.co/datasets/DimmsMath/Indonesia-Stock-Symbols-and-Metadata.textn<1K0 likes48 downloads4mo agoHugging Face14DGurgurov /indonesian_sa Sentiment Analysis Data for the Indonesian Language Dataset Description: This dataset contains a sentiment analysis data from Purwarianti et al. (2019). Data Structure: The data was used for the project on injecting external commonsense knowledge into multilingual Large Language Models. Citation: @inproceedings{purwarianti2019improving, title={Improving bi-lstm performance for indonesian sentiment analysis using paragraph vector}, author={Purwarianti, Ayu and Crisdayanti, Ida… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/indonesian_sa.texttext-classification10K<n<100K1 likes47 downloads2y agoHugging Face15Ichsan2895 /OASST_Top1_IndonesianBase dataset : OpenAssistant/oasst1 We selected the dataset with "english" language & has rank = 1st. Finally, we translate it to Indonesian with Marian NMT and pretrained model from Helsinki-NLP/opus-mt-en-id. CITATION @InProceedings{mariannmt, title = {Marian: Fast Neural Machine Translation in {C++}}, author = {Junczys-Dowmunt, Marcin and Grundkiewicz, Roman and Dwojak, Tomasz and Hoang, Hieu and Heafield, Kenneth and Neckermann… See the full description on the dataset page: https://huggingface.co/datasets/Ichsan2895/OASST_Top1_Indonesian.textquestion-answering1K<n<10K6 likes40 downloads3y agoHugging Face16glhpradipta /alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian We wrangled the original dataset format to 'input' & 'output' format. For example: BEFORE: [ { "from": "human", "value": "Saranlah slogan untuk kampanye daur ulang\n" }, { "from": "gpt", "value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \ "Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \ "Daur… See the full description on the dataset page: https://huggingface.co/datasets/glhpradipta/alpaca-gpt4-indonesian.textquestion-answering10K<n<100K0 likes40 downloads14d agoHugging Face17fadhilakbar /indonesian-names-datasettext10K<n<100K0 likes38 downloads4mo agoHugging Face18egdrga /indonesian-twitter-hate-speech-cleaned Dataset Card for indonesian-twitter-hate-speech-cleaned Dataset Summary Cleaned Indonesian Twitter Hate Speech is a curated dataset consisting of Indonesian-language tweets labeled as either hate or neutral. The dataset was collected through a combination of direct scraping from Twitter and aggregation from multiple publicly available GitHub repositories. The data has been cleaned to remove duplicates, irrelevant content, and non-textual noise, making it suitable… See the full description on the dataset page: https://huggingface.co/datasets/egdrga/indonesian-twitter-hate-speech-cleaned.texttext-classification10K<n<100K0 likes38 downloads8d agoHugging Face19Speech-data /Indonesian-Speech-Dataset 🎧 Indonesian Speech Dataset The Indonesian Speech Dataset is a high-quality speech audio dataset designed to deliver structured and scalable audio data for AI-powered voice systems. It contains 162 hours of audio data across 821 files, provided in MP3 and WAV formats, with a total size of 210 MB. This well-curated audio dataset ensures balanced and representative voice data, with 51% female and 49% male speakers, and a broad age distribution from 18 to 50+ years. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Indonesian-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes37 downloads6mo agoHugging Face20abid /indonesia-medical-qna indonesia-medical-qna This dataset is configured so the Hugging Face Dataset Viewer loads qna.csv as the primary data file. The repository also contains all_links.csv, which has a different schema and is kept as an auxiliary file rather than part of the default viewer configuration. This avoids schema-casting errors caused by the Hub trying to combine both CSV files into one split. text100K<n<1M5 likes34 downloads6mo agoHugging Face21casecrit /2024-indonesian-electionThe dataset encompasses news articles spanning from November 29, 2023, to February 6, 2024, capturing the discourse surrounding the five presidential debates orchestrated by the General Elections Commission. Sourced from reputable platforms such as detik, kompas, and liputan6, the dataset offers a comprehensive insight into the electoral landscape and the media coverage thereof. text10K<n<100K4 likes33 downloads3y agoHugging Face22manueltonneau /indonesian-hate-speech-supersetgated Indonesian Hate Speech Superset This dataset is a superset (N=14,306) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Indonesian hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that: are documented are publicly available focus on hate speech, defined broadly as "any kind of communication in speech, writing or… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/indonesian-hate-speech-superset.tabulartext-classification10K<n<100K4 likes31 downloads2y agoHugging Face23talithaolga /Indonesian-Emotion-Classificationtexttext-classification10K<n<100K4 likes31 downloads1y agoHugging Face24RimuruChan67 /indonesian-googleplay-sentiment Indonesian Google Play Review Sentiment Dataset A sentiment classification dataset of 5,336 Indonesian-language Google Play reviews collected from three widely used Indonesian applications: Gojek (on-demand transport and delivery), Tokopedia, and Shopee (e-commerce). The dataset accompanies the paper Sentiment Classification of Indonesian Google Play Reviews: A Hybrid Evaluation of Random Forest, CNN, and IndoBERT (ICORIS 2026). Collection Reviews were retrieved… See the full description on the dataset page: https://huggingface.co/datasets/RimuruChan67/indonesian-googleplay-sentiment.texttext-classification1K<n<10K0 likes29 downloads1mo agoHugging Face25pauwdanny /indonesian_hoax_news_oritextn<1K2 likes28 downloads4y agoHugging Face26Chemin-AI /Indonesian_DatasetThis dataset is a collection of 50 questions that consists of 4 categories: language, domain, geographical, and combined.Each question has two variations: English & Indonesian. Statistics: Language-Based: 15 Domain-Based: 15 Geographical-Based: 15 Combined: 5 texttranslationn<1K2 likes27 downloads2y agoHugging Face27DGurgurov /indonesian_conceptnet ConceptNet Data for the Indonesian Language Dataset Description: This dataset contains data extracted from ConceptNet using the dedicated module for fetching knowledge from the graph, available on GitHub. Data Structure: The data is converted from triplets into natural text using a pre-defined relationship mapping and split into training and validation sets. It was used for training language adapters for the project aimed at injecting external commonsense knowledge into multilingual… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/indonesian_conceptnet.text10K<n<100K0 likes23 downloads2y agoHugging Face28web3hungry /indonesia-affordable-housing Indonesia Affordable Housing Dataset A comprehensive dataset of affordable housing projects across Indonesia, containing detailed information about residential properties, specifications, locations, pricing, and developer data. Dataset Creator: web3hungry Dataset ID: web3hungry/indonesia-affordable-housing License: CC0 1.0 Dataset Overview This dataset provides extensive data on housing developments throughout Indonesia, covering both subsidized and commercial… See the full description on the dataset page: https://huggingface.co/datasets/web3hungry/indonesia-affordable-housing.tabulartabular-classification10K<n<100K0 likes23 downloads7mo agoHugging Face29Hadisawara /indonesian-regional-tax-revenue Indonesian Regional Tax Revenue Analytics Dataset Dataset Description Dataset ini berisi data sintetis realisasi penerimaan pajak daerah di Indonesia, mencakup 10 provinsi, 100+ kota/kabupaten, 13 jenis pajak daerah, dan rentang waktu 7 tahun (2018–2024). Dataset dirancang untuk mendukung penelitian analitik, forecasting, dan klasifikasi di bidang kebijakan fiskal dan tata kelola keuangan daerah. Struktur data terinspirasi dari sistem SIMPAD (Sistem Informasi… See the full description on the dataset page: https://huggingface.co/datasets/Hadisawara/indonesian-regional-tax-revenue.tabulartabular-classification10K<n<100K0 likes22 downloads3mo agoHugging Face30pauwdanny /indonesian_hoax_news_datasettextn<1K1 likes20 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.