datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
public-javanese-dataset
Public-Domain Javanese Manuscript & Text Dataset
A curated collection of public-domain (or openly-licensed) Javanese-language
source material — manuscript scans, plain-text transcriptions, digitized
printed books, and aksara Jawa (Carakan) primers.
What's in it
#
Directory
Title
Material
Author / Credit
License
1
kakawin-nagarakertagama
Kakawin Nagarakertagama (Desawarnnana)
Old Javanese (Kawi) kakawin
Mpu Prapanca (1365)
Public domain
2… See the full description on the dataset page: https://huggingface.co/datasets/thesimonharms/public-javanese-dataset.SLR35_javaneseimdb-javaneseLarge Movie Review Dataset translated to Javanese.
This is a dataset for binary sentiment classification containing substantially
more data than previous benchmark datasets. We provide a set of 25,000 highly
polar movie reviews for training, and 25,000 for testing. There is additional
unlabeled data for use as well. We translated the original IMDB Dataset to
Javanese using the multi-lingual MarianMT Transformer model from
`Helsinki-NLP/opus-mt-en-mul`.javanese-Komodo-pixelgpt
Javanese PixelGPT Dataset
This dataset contains preprocessed Javanese text data for training PixelGPT models.
Dataset Statistics
Language: Javanese (jawa)
Total samples: 401,542
Train samples: 400,726
Test samples: 816
Tokenizers
[TO BE EDITED]
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel representation of aksara text
<tokenizer_name>_token_ids: Token IDs from grapheme-based… See the full description on the dataset page: https://huggingface.co/datasets/Exqrch/javanese-Komodo-pixelgpt.Rebuttal-javanese-pixelgpt
Javanese PixelGPT Tokenizer Ablation Dataset
Optimized with Font Size 6 and Dynamic Trimming.
Tokenizer Schema
tok_grapheme: Language-specific Grapheme BPE (izzako/javanese-llama-tokenizer)
tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT)
tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1)
tok_mt5: Google Multilingual Unigram (google/mt5-small)
alpaca-javanese-cleanedThis repository contains the dataset used for the TaCo paper.
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation
@inproceedings{upadhayay2024taco,
title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes},
author={Bibek Upadhayay and Vahid Behzadan},
booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-javanese-cleaned.Gatra-1-Javanese
GatraOne (Gatra-1) is a synthethic Jawa Krama instruction-tuning dataset, generated by GPT-4.
Introducing the Gatra-1 dataset
This is a synthetic dataset to fine-tune LLMs into responding in Jawa Krama, the high-register of Javanese language. It is 98% generated using GPT-4, which has very good Jawa Krama capabilities. It is currently a 'beta' version with only 560 input-output prompts.
So far, this has been only tested on fine-tuning GPT-3.5 with considerable success.… See the full description on the dataset page: https://huggingface.co/datasets/afrizalha/Gatra-1-Javanese.Unggah-Ungguh
Javanese Honorifics Dataset (Unggah-Ungguh - Released Version)
The Javanese language, spoken by over 98 million people, features a distinctive honorific system known as Unggah-Ungguh Basa. In this dataset we present UNGGAH-UNGGUH, a carefully curated dataset designed to encapsulate the nuances of Unggah-Ungguh Basa, the Javanese speech etiquette framework that dictates the choice of words and phrases based on social hierarchy and context.
Paper: https://arxiv.org/pdf/2502.20864… See the full description on the dataset page: https://huggingface.co/datasets/JavaneseHonorifics/Unggah-Ungguh.javanese_sundanese_story_clozejavanese-speech-datasetjavanesejavanese_sa
Sentiment Analysis Data for the Javanese Language
Dataset Description:
This dataset contains a sentiment analysis data from Wongso et al. (2021).
Data Structure:
The data was used for the project on injecting external commonsense knowledge into multilingual Large Language Models.
Citation:
@inproceedings{wongso2021causal,
title={Causal and Masked Language Modeling of Javanese Language using Transformer-based Architectures},
author={Wongso, Wilson and Setiawan, David Samuel… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/javanese_sa.alpaca_javanese_tacoThis repository contains the dataset used for the TaCo paper.
The dataset follows the style outlined in the TaCo paper, as follows:
{
"instruction": "instruction in xx",
"input": "input in xx",
"output": "Instruction in English: instruction in en ,
Response in English: response in en ,
Response in xx: response in xx "
}
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_javanese_taco.javanese_conceptnet
ConceptNet Data for the Javanese Language
Dataset Description:
This dataset contains data extracted from ConceptNet using the dedicated module for fetching knowledge from the graph, available on GitHub.
Data Structure:
The data is converted from triplets into natural text using a pre-defined relationship mapping and split into training and validation sets. It was used for training language adapters for the project aimed at injecting external commonsense knowledge into multilingual… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/javanese_conceptnet.javanese-to-indonesiaJavaneseIMDBClassification
JavaneseIMDBClassification
An MTEB dataset
Massive Text Embedding Benchmark
Large Movie Review Dataset translated to Javanese. This is a dataset for binary sentiment classification containing substantially more data than previous benchmark datasets.
Task category
t2c
Domains
Reviews, Written
Reference
https://github.com/w11wo/nlp-datasets#javanese-imdb
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following… See the full description on the dataset page: https://huggingface.co/datasets/mteb/JavaneseIMDBClassification.Centhini-1-Javanese
Dataset details
The dataset comprises 529,575 pretraining examples for both Ngoko and Krama Javanese. The data is almost predominantly translation generated with Deepseek V3. The translations include data from English language Fineweb and a paraphrased translation from Indonesian mc4 dataset. Other examples here include ancient Javanese texts, like Serat Centhini and Babad Tanah Djawi, but also open texts like Javanese wikipedia.
To our knowledge, this is the largest easily… See the full description on the dataset page: https://huggingface.co/datasets/afrizalha/Centhini-1-Javanese.javanese-translatedGatra-2-Javanese
Dataset details
The dataset comprises 36870 prompt-response pairs of Krama Javanese instruction-tuning examples. The data is almost entirely synthetic with minimal human curation. The current dataset supports only single-turn QA, although fine-tuning on instruction-tuned models may allow for transfer of multi-turn capabilities.
The prompts are generated by GPT-4o, while the responses are generated by Claude 3 Haiku. The way the data set was generated, the prompt may contain terms in… See the full description on the dataset page: https://huggingface.co/datasets/afrizalha/Gatra-2-Javanese.alpaca-javanese-7k-cleanedJavanese-Speech-Dataset
🎧 Javanese Speech Dataset
The Javanese Speech Dataset is a structured and scalable speech audio dataset designed to provide high-quality audio data for training modern AI and machine learning models. It includes 85 hours of audio data across 585 files, delivered in MP3 and WAV formats, with a total size of 104 MB. This well-balanced audio dataset offers diverse and representative voice data, with 51% female and 49% male speakers, and an age range spanning from 18 to 50+ years. The… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Javanese-Speech-Dataset.oasst-javanese
Dataset Summary
We translated the OpenAssistant Conversations (OASST) dataset into Javanese using Meta's No Language Left Behind (NLLB) model.
Why Javanese?
Javanese is spoken by over 90 million people on the island of Java in Indonesia. While its prevalence is comparable to other widely spoken languages, such as Vietnamese and Turkish, its representation in current large language model (LLM) chatbots remains limited. By translating this dataset, we aim to enhance the… See the full description on the dataset page: https://huggingface.co/datasets/richardcsuwandi/oasst-javanese.fork-google-openslr-javanesejavanese-multilingual-lexicon
Javanese Multilingual Lexicon
A multilingual lexicon dataset of 202,912 entries derived from Sastra.org, a digital archive of Javanese literary heritage maintained by Yayasan Sastra Lestari.
The dataset compiles 32 historical lexicographic sources spanning from the 1830s to 2010, covering Javanese, Old Javanese (Kawi), Indonesian, Dutch, English, and French.
Dataset Structure
The dataset is provided as JSONL files. The full corpus is in data/leksikon_all.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/junwatu/javanese-multilingual-lexicon.Javanese-Non-STEM-Textbook-DatasetDataset Description:
This dataset is a large-scale collection of Javanese Non-STEM textbook data, containing 597 books and 67.05 million words, designed to support the development and training of advanced NLP systems and AI models for language understanding, reasoning, and general knowledge learning in Javanese.
Full Dataset Overview
This dataset is part of a large-scale multilingual educational corpus containing over 3+ billion words across 5,000+ subjects, supported by interwoven images for… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Javanese-Non-STEM-Textbook-Dataset.ms-javanesealpaca-clean-javanesejavanese-pixelgpt
Javanese PixelGPT Dataset
This dataset contains preprocessed Javanese text data for training PixelGPT models.
Dataset Statistics
Language: Javanese (jawa)
Total samples: 401,542
Train samples: 400,726
Test samples: 816
Tokenizers
Grapheme tokenizer: izzako/javanese-llama-tokenizer
LLaMA tokenizer: ernie-research/DualGPT
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel… See the full description on the dataset page: https://huggingface.co/datasets/izzako/javanese-pixelgpt.javanese-collectionjavanese_asr_dataset_20k
