datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
quran-terjemahan-indonesia-audio-by-ayah
Audio Terjemahan Al-Qur'an Indonesia per Ayat
Dataset ini berisi audio terjemahan Al-Qur'an bahasa Indonesia yang dipisahkan per ayat.
Audio dibuat menggunakan Gemini TTS dari teks terjemahan Al-Qur'an bahasa Indonesia. Setiap file audio mewakili satu ayat.
Isi Dataset
6.236 file audio WAV
6.236 file audio M4A/AAC terkompresi untuk aplikasi mobile
Bahasa Indonesia
Satu file audio untuk setiap ayat
Format WAV dan M4A/AAC
Disusun berdasarkan nomor surah dan nomor… See the full description on the dataset page: https://huggingface.co/datasets/insanalamin/quran-terjemahan-indonesia-audio-by-ayah.alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian
We wrangled the original dataset format to 'input' & 'output' format. For example:
BEFORE:
[ { "from": "human",
"value": "Saranlah slogan untuk kampanye daur ulang\n" },
{ "from": "gpt",
"value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \
"Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \
"Daur… See the full description on the dataset page: https://huggingface.co/datasets/Ichsan2895/alpaca-gpt4-indonesian.mc4-idA thoroughly cleaned version of the Italian portion of the multilingual
colossal, cleaned version of Common Crawl's web crawl corpus (mC4) by AllenAI.
Based on Common Crawl dataset: "https://commoncrawl.org".
This is the processed version of Google's mC4 dataset by AllenAI, with further cleaning
detailed in the repository README file.Indonesia-Stock-Symbols-and-Metadata
Indonesia Stock Symbols & Company Metadata
This dataset contains stock symbols and basic company metadata for all listed companies in Indonesia.It is updated weekly if new changes are there.
📊 Dataset Contents
The dataset is provided as a CSV file with the following columns:
Column
Description
name
Full company name
ticker
Stock ticker symbol (e.g., AAPL, MSFT)
market
The exchange/market where the stock is listed
sector
The primary business sector of… See the full description on the dataset page: https://huggingface.co/datasets/kjhq/Indonesia-Stock-Symbols-and-Metadata.t5gemma2-indonesia-instruct-v1
T5Gemma-2 Indonesian Instruct — Mono-Repo
Satu repositori dataset HF untuk seluruh data pelatihan T5-Gemma-2 bahasa Indonesia.
Diorganisasi per fungsi (fondasi → spesifik → preferensi) dengan folder/subfolder,
setiap config = folder dan berisi split train + validation (80:20) di level percakapan.
Struktur (by fungsi)
t5gemma2-indonesia-instruct-v1/
├── README.md
├── manifest.json
├── chat_idx_map.json
├── foundation/ ← FASE 1 · fondasi Bahasa… See the full description on the dataset page: https://huggingface.co/datasets/daruokta/t5gemma2-indonesia-instruct-v1.Indonesian_Regulation_QA
📘 Indonesian Regulation QA Dataset (From Public Regulation Sources)
This dataset consists of automatically generated legal question-answer pairs, where questions are crafted from basic legal inquiries, and answers are mapped directly to parsed Indonesian regulation articles from alternative public legal repositories.
📌 Dataset Overview
Source: Public Indonesian regulation databases and portals
QA Generation:
Basic legal questions (template-driven or commonly asked)… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/Indonesian_Regulation_QA.indonesian-voice-transcription-1.4.9a-rindonesian-voice-transcription-1.4.9a.2indonesian-id-card-dummy
Indonesian KTP Dataset 24K (Flat & Augmented - Commercially Safe)
Welcome to the Indonesian KTP (Kartu Tanda Penduduk) Dataset. This is a highly robust, high-fidelity, and commercially safe synthetic dataset designed to advance SOTA (State-of-the-Art) research in Document Information Extraction (DIE), Key Information Extraction (KIE), and Optical Character Recognition (OCR) specifically for Indonesian National ID cards (KTP-el).
The dataset is natively packaged in Apache Parquet… See the full description on the dataset page: https://huggingface.co/datasets/cloverx-id/indonesian-id-card-dummy.cleaned-data-split-0
Dataset Card for "cleaned-data-split-0"
More Information needed
Indonesian-ASR-11-Class-Dataset
Indonesian ASR 11-Class Dataset
Public Hugging Face repository for an Indonesian ASR corpus and its paper-supporting benchmark artifacts.
Dataset summary
Audio files: 104,500 WAV files
Real/human recordings: 104,368
Synthetic repair files: 132
Sentence classes: 11 Indonesian sentence categories
Canonical balanced sentence slots: 209 (11 categories × 19 retained slots)
Public speaker labels: M1..M12, F1..F8, plus synthetic labels Ms*/Fs*
Audio format: 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/Atika88/Indonesian-ASR-11-Class-Dataset.pmd_indonesiaalpaca-gpt4-indonesianThe dataset is used in the research related to MultilingualSIFT.
indonesian-voice-transcription-1.4.85aindonesian-voice-transcription-1.4.9a-cv-fl-slrjv-mdindonesian-voice-transcription-1.3.9cindonesian-voice-transcription-1.4.9rcvindonesian-voice-transcription-1.4.9awikipedia-idindonesian-voice-transcription-1.5.9a-cv-fl-slrjv-mdwikipedia-10kindonesia-slangDataset-Text-To-Speech-Indonesia
🎵 Dataset Audio Bahasa Indonesia
Dataset audio berkualitas tinggi untuk Text-to-Speech (TTS) bahasa Indonesia.
Dibuat oleh : Muhammad Arief, S.Kom.Universitas Muhammadiyah SorongTeknik Informatika 2020
📊 Spesifikasi Teknis
Parameter
Nilai
Satuan
Total Durasi
16.38
jam
Jumlah Segmen
4531
file
Durasi Rata-rata
13.01
detik
Sample Rate KHz
22
kHz
Sample Rate Hz
22000
Hz
Bit Depth
PCM_16
PCM
Format
wav
Lossless
🔄 Urutan Pengolahan… See the full description on the dataset page: https://huggingface.co/datasets/X-lord/Dataset-Text-To-Speech-Indonesia.indonesian-voice-transcription-1.4.9rtwitter_indonesia_sarcastic
Twitter Indonesia Sarcastic
Twitter Indonesia Sarcastic is a dataset intended for sarcasm detection in the Indonesian language. This dataset is introduced in Khotijah et al. (2020), whereby Indonesian tweets are collected and labeled as either sarcastic or non-sarcastic. We took the raw data, and performed several cleaning procedures such as: sentence order re-reversal, deduplication with minHash LSH, PII masking to remove usernames, hashtags, emails, URLs, and finally a random… See the full description on the dataset page: https://huggingface.co/datasets/w11wo/twitter_indonesia_sarcastic.indonesian-voice-transcription-1.0e-commerce-sentiment-bahasa-indonesia
E-Commerce Sentiment Analysis Dataset (Indonesian)
Dataset komentar dan ulasan produk e-commerce dalam Bahasa Indonesia untuk analisis sentiment.
Dataset Summary
Dataset ini berisi 21,840 komentar e-commerce dalam Bahasa Indonesia yang telah dilabeli dengan sentiment (positif, netral, negatif). Dataset mencakup berbagai jenis komentar termasuk sarkasme dan ironi yang umum ditemukan dalam ulasan online.
Dataset Structure
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/Azzamwsa233/e-commerce-sentiment-bahasa-indonesia.Indonesian_Regulation_QA
📘 Indonesian Regulation QA Dataset (From Public Regulation Sources)
This dataset consists of automatically generated legal question-answer pairs, where questions are crafted from basic legal inquiries, and answers are mapped directly to parsed Indonesian regulation articles from alternative public legal repositories.
📌 Dataset Overview
Source: Public Indonesian regulation databases and portals
QA Generation:
Basic legal questions (template-driven or commonly asked)… See the full description on the dataset page: https://huggingface.co/datasets/horelulus/Indonesian_Regulation_QA.IndonesianIdClickbaitClassification
IndonesianIdClickbaitClassification
An MTEB dataset
Massive Text Embedding Benchmark
The CLICK-ID dataset is a collection of Indonesian news headlines that was collected from 12 local online news publishers.
Task category
t2c
Domains
News, Written
Reference
http://www.sciencedirect.com/science/article/pii/S2352340920311252
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/IndonesianIdClickbaitClassification.bahasa-indonesia-corpus
