datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
balinese-Komodo-pixelgpt
Balinese PixelGPT Dataset
This dataset contains preprocessed Balinese text data for training PixelGPT models.
Dataset Statistics
Language: Balinese (bali)
Total samples: 54,467
Train samples: 54,017
Test samples: 450
Tokenizers
[TO BE EDITED]
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel representation of aksara text
<tokenizer_name>_token_ids: Token IDs from grapheme-based… See the full description on the dataset page: https://huggingface.co/datasets/Exqrch/balinese-Komodo-pixelgpt.balinese-pixelgpt
Balinese PixelGPT Dataset
This dataset contains preprocessed Balinese text data for training PixelGPT models.
Dataset Statistics
Language: Balinese (bali)
Total samples: 54,467
Train samples: 54,017
Test samples: 450
Tokenizers
Grapheme tokenizer: izzako/javanese-llama-tokenizer
LLaMA tokenizer: ernie-research/DualGPT
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel… See the full description on the dataset page: https://huggingface.co/datasets/izzako/balinese-pixelgpt.balinese-cpt-corpus
Balinese continued-pretraining corpus (v0)
Cleaned monolingual Balinese for LLM continued-pretraining (FineWeb-2 ban_Latn + Balinese Wikipedia + story/text sets; normalize + quality-filter + exact-dedup). Part of Open Indonesia Models.
{
"docs": 29825,
"stats": {
"in": 33670,
"empty": 4,
"rejected": 3539,
"dup": 302,
"kept": 29825
},
"per_source": {
"HuggingFaceFW/fineweb-2": 9233,
"wikimedia/wikipedia": 20986… See the full description on the dataset page: https://huggingface.co/datasets/timothydillan/balinese-cpt-corpus.
