datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
javanese-Komodo-pixelgpt
Javanese PixelGPT Dataset
This dataset contains preprocessed Javanese text data for training PixelGPT models.
Dataset Statistics
Language: Javanese (jawa)
Total samples: 401,542
Train samples: 400,726
Test samples: 816
Tokenizers
[TO BE EDITED]
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel representation of aksara text
<tokenizer_name>_token_ids: Token IDs from grapheme-based… See the full description on the dataset page: https://huggingface.co/datasets/Exqrch/javanese-Komodo-pixelgpt.Gatra-1-Javanese
GatraOne (Gatra-1) is a synthethic Jawa Krama instruction-tuning dataset, generated by GPT-4.
Introducing the Gatra-1 dataset
This is a synthetic dataset to fine-tune LLMs into responding in Jawa Krama, the high-register of Javanese language. It is 98% generated using GPT-4, which has very good Jawa Krama capabilities. It is currently a 'beta' version with only 560 input-output prompts.
So far, this has been only tested on fine-tuning GPT-3.5 with considerable success.… See the full description on the dataset page: https://huggingface.co/datasets/afrizalha/Gatra-1-Javanese.Centhini-1-Javanese
Dataset details
The dataset comprises 529,575 pretraining examples for both Ngoko and Krama Javanese. The data is almost predominantly translation generated with Deepseek V3. The translations include data from English language Fineweb and a paraphrased translation from Indonesian mc4 dataset. Other examples here include ancient Javanese texts, like Serat Centhini and Babad Tanah Djawi, but also open texts like Javanese wikipedia.
To our knowledge, this is the largest easily… See the full description on the dataset page: https://huggingface.co/datasets/afrizalha/Centhini-1-Javanese.Gatra-2-Javanese
Dataset details
The dataset comprises 36870 prompt-response pairs of Krama Javanese instruction-tuning examples. The data is almost entirely synthetic with minimal human curation. The current dataset supports only single-turn QA, although fine-tuning on instruction-tuned models may allow for transfer of multi-turn capabilities.
The prompts are generated by GPT-4o, while the responses are generated by Claude 3 Haiku. The way the data set was generated, the prompt may contain terms in… See the full description on the dataset page: https://huggingface.co/datasets/afrizalha/Gatra-2-Javanese.javanese-multilingual-lexicon
Javanese Multilingual Lexicon
A multilingual lexicon dataset of 202,912 entries derived from Sastra.org, a digital archive of Javanese literary heritage maintained by Yayasan Sastra Lestari.
The dataset compiles 32 historical lexicographic sources spanning from the 1830s to 2010, covering Javanese, Old Javanese (Kawi), Indonesian, Dutch, English, and French.
Dataset Structure
The dataset is provided as JSONL files. The full corpus is in data/leksikon_all.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/junwatu/javanese-multilingual-lexicon.javanese-pixelgpt
Javanese PixelGPT Dataset
This dataset contains preprocessed Javanese text data for training PixelGPT models.
Dataset Statistics
Language: Javanese (jawa)
Total samples: 401,542
Train samples: 400,726
Test samples: 816
Tokenizers
Grapheme tokenizer: izzako/javanese-llama-tokenizer
LLaMA tokenizer: ernie-research/DualGPT
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel… See the full description on the dataset page: https://huggingface.co/datasets/izzako/javanese-pixelgpt.javanese_instruct_storiesA dataset of parallel translation-based instructions for Javanese language as a target language.
Materials are taken from randomly selected children stories at https://storyweaver.org.in, under CC-By-SA-4.0 license.
The template IDs are:
(1, 'Terjemahno penggalan teks crito ing ngisor iki saka Bahasa Inggris dadi teks crito ing Basa Jawa:', 'Terjemahane utawa padanan teks crito kasebut ing Basa Jawa yaiku:'),
(2, 'Terjemahno penggalan teks crito ing ngisor iki saka Bahasa Indonesia dadi teks… See the full description on the dataset page: https://huggingface.co/datasets/Iftitahu/javanese_instruct_stories.sea-javanese-cleaned-parquet-v1
SEA Javanese Cleaned Parquet v1
Dataset Summary
This dataset is a cleaned Javanese pretraining corpus exported in Hugging Face parquet format.
Current public sources used in this release:
HuggingFaceFW/fineweb-2 / jav_Latn
allenai/c4 / jv
afrizalha/Centhini-1-Javanese
Cleaning and Deduplication
Current pipeline:
basic text cleaning
short-text filtering
repetition filtering
rule-based noise filtering
document-level exact deduplication across all included… See the full description on the dataset page: https://huggingface.co/datasets/rinnieyoung/sea-javanese-cleaned-parquet-v1.
