CoolFace
Datasetpublic

Exqrch/balinese-Komodo-pixelgpt

Balinese PixelGPT Dataset This dataset contains preprocessed Balinese text data for training PixelGPT models. Dataset Statistics Language: Balinese (bali) Total samples: 54,467 Train samples: 54,017 Test samples: 450 Tokenizers [TO BE EDITED] Features text_id: Document identifier chunk_id: Chunk identifier within document pixel_values: Rendered pixel representation of aksara text <tokenizer_name>_token_ids: Token IDs from… See the full description on the dataset page: https://huggingface.co/datasets/Exqrch/balinese-Komodo-pixelgpt.

sourceHugging Facecc-by-4.0updated 7mo agoView on Hugging Face
0likes73downloads

Exqrch/balinese-Komodo-pixelgpt · main · files are served by the source, never re-hosted here