Javanese
Datasets
All datasets matching “Javanese”public-javanese-dataset
Public-Domain Javanese Manuscript & Text Dataset
A curated collection of public-domain (or openly-licensed) Javanese-language
source material — manuscript scans, plain-text transcriptions, digitized
printed books, and aksara Jawa (Carakan) primers.
What's in it
#
Directory
Title
Material
Author / Credit
License
1
kakawin-nagarakertagama
Kakawin Nagarakertagama (Desawarnnana)
Old Javanese (Kawi) kakawin
Mpu Prapanca (1365)
Public domain
2… See the full description on the dataset page: https://huggingface.co/datasets/thesimonharms/public-javanese-dataset.SLR35_javaneseimdb-javaneseLarge Movie Review Dataset translated to Javanese.
This is a dataset for binary sentiment classification containing substantially
more data than previous benchmark datasets. We provide a set of 25,000 highly
polar movie reviews for training, and 25,000 for testing. There is additional
unlabeled data for use as well. We translated the original IMDB Dataset to
Javanese using the multi-lingual MarianMT Transformer model from
`Helsinki-NLP/opus-mt-en-mul`.javanese-Komodo-pixelgpt
Javanese PixelGPT Dataset
This dataset contains preprocessed Javanese text data for training PixelGPT models.
Dataset Statistics
Language: Javanese (jawa)
Total samples: 401,542
Train samples: 400,726
Test samples: 816
Tokenizers
[TO BE EDITED]
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel representation of aksara text
<tokenizer_name>_token_ids: Token IDs from grapheme-based… See the full description on the dataset page: https://huggingface.co/datasets/Exqrch/javanese-Komodo-pixelgpt.Rebuttal-javanese-pixelgpt
Javanese PixelGPT Tokenizer Ablation Dataset
Optimized with Font Size 6 and Dynamic Trimming.
Tokenizer Schema
tok_grapheme: Language-specific Grapheme BPE (izzako/javanese-llama-tokenizer)
tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT)
tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1)
tok_mt5: Google Multilingual Unigram (google/mt5-small)
alpaca-javanese-cleanedThis repository contains the dataset used for the TaCo paper.
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation
@inproceedings{upadhayay2024taco,
title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes},
author={Bibek Upadhayay and Vahid Behzadan},
booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-javanese-cleaned.
