datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
javanese-Komodo-pixelgpt
Javanese PixelGPT Dataset
This dataset contains preprocessed Javanese text data for training PixelGPT models.
Dataset Statistics
Language: Javanese (jawa)
Total samples: 401,542
Train samples: 400,726
Test samples: 816
Tokenizers
[TO BE EDITED]
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel representation of aksara text
<tokenizer_name>_token_ids: Token IDs from grapheme-based… See the full description on the dataset page: https://huggingface.co/datasets/Exqrch/javanese-Komodo-pixelgpt.balinese-Komodo-pixelgpt
Balinese PixelGPT Dataset
This dataset contains preprocessed Balinese text data for training PixelGPT models.
Dataset Statistics
Language: Balinese (bali)
Total samples: 54,467
Train samples: 54,017
Test samples: 450
Tokenizers
[TO BE EDITED]
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel representation of aksara text
<tokenizer_name>_token_ids: Token IDs from grapheme-based… See the full description on the dataset page: https://huggingface.co/datasets/Exqrch/balinese-Komodo-pixelgpt.sundanese-pixelgpt
Sundanese PixelGPT Dataset
This dataset contains preprocessed Sundanese text data for training PixelGPT models.
Dataset Statistics
Language: Sundanese (sunda)
Total samples: 294,756
Train samples: 293,933
Test samples: 823
Tokenizers
Grapheme tokenizer: izzako/sunda-llama-tokenizer
LLaMA tokenizer: ernie-research/DualGPT
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel representation… See the full description on the dataset page: https://huggingface.co/datasets/izzako/sundanese-pixelgpt.balinese-pixelgpt
Balinese PixelGPT Dataset
This dataset contains preprocessed Balinese text data for training PixelGPT models.
Dataset Statistics
Language: Balinese (bali)
Total samples: 54,467
Train samples: 54,017
Test samples: 450
Tokenizers
Grapheme tokenizer: izzako/javanese-llama-tokenizer
LLaMA tokenizer: ernie-research/DualGPT
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel… See the full description on the dataset page: https://huggingface.co/datasets/izzako/balinese-pixelgpt.lampung-pixelgpt
Lampung PixelGPT Dataset
This dataset contains preprocessed Lampung text data for training PixelGPT models.
Dataset Statistics
Language: Lampung (lampung)
Total samples: 1,029
Train samples: 945
Test samples: 84
Tokenizers
Grapheme tokenizer: izzako/sunda-llama-tokenizer
LLaMA tokenizer: ernie-research/DualGPT
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel representation of aksara… See the full description on the dataset page: https://huggingface.co/datasets/izzako/lampung-pixelgpt.javanese-pixelgpt
Javanese PixelGPT Dataset
This dataset contains preprocessed Javanese text data for training PixelGPT models.
Dataset Statistics
Language: Javanese (jawa)
Total samples: 401,542
Train samples: 400,726
Test samples: 816
Tokenizers
Grapheme tokenizer: izzako/javanese-llama-tokenizer
LLaMA tokenizer: ernie-research/DualGPT
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel… See the full description on the dataset page: https://huggingface.co/datasets/izzako/javanese-pixelgpt.
