CoolFace
Datasetpublic

izzako/javanese-pixelgpt

Javanese PixelGPT Dataset This dataset contains preprocessed Javanese text data for training PixelGPT models. Dataset Statistics Language: Javanese (jawa) Total samples: 401,542 Train samples: 400,726 Test samples: 816 Tokenizers Grapheme tokenizer: izzako/javanese-llama-tokenizer LLaMA tokenizer: ernie-research/DualGPT Features text_id: Document identifier chunk_id: Chunk identifier within document pixel_values: Rendered pixel… See the full description on the dataset page: https://huggingface.co/datasets/izzako/javanese-pixelgpt.

sourceHugging Facecc-by-4.0updated 9mo agoView on Hugging Face
0likes11downloads
settings

This repository belongs to izzako on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namejavanese-pixelgpt
visibilitypublic
licencecc-by-4.0
gatedno
ownerizzako
Account settings
izzako/javanese-pixelgpt · CoolFace