CoolFace
10 results

pixelgpt

unstonio /pixelgpt-24x24-20k PixelGPT 24×24 — 20K 20,000 native 24×24 pixel-art sprites with captions and semantic taxonomy labels. This is a clean, rebalanced, rights-conscious public subset of the larger PixelGPT 24×24 dataset. Every sprite: is rendered at a native resolution of 24×24 pixels uses no more than 5 colors includes an original text caption is assigned to a two-level semantic taxonomy is distributed in lossless PNG and Parquet formats Looking for the complete dataset?The full edition… See the full description on the dataset page: https://huggingface.co/datasets/unstonio/pixelgpt-24x24-20k.imagetext-to-image10K<n<100K69 likes538 downloads2mo agoHugging FaceExqrch /javanese-Komodo-pixelgpt Javanese PixelGPT Dataset This dataset contains preprocessed Javanese text data for training PixelGPT models. Dataset Statistics Language: Javanese (jawa) Total samples: 401,542 Train samples: 400,726 Test samples: 816 Tokenizers [TO BE EDITED] Features text_id: Document identifier chunk_id: Chunk identifier within document pixel_values: Rendered pixel representation of aksara text <tokenizer_name>_token_ids: Token IDs from grapheme-based… See the full description on the dataset page: https://huggingface.co/datasets/Exqrch/javanese-Komodo-pixelgpt.tabulartext-generation100K<n<1M0 likes100 downloads7mo agoHugging FaceExqrch /Rebuttal-javanese-pixelgpt Javanese PixelGPT Tokenizer Ablation Dataset Optimized with Font Size 6 and Dynamic Trimming. Tokenizer Schema tok_grapheme: Language-specific Grapheme BPE (izzako/javanese-llama-tokenizer) tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT) tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1) tok_mt5: Google Multilingual Unigram (google/mt5-small) tabular100K<n<1M0 likes83 downloads6mo agoHugging FaceExqrch /balinese-Komodo-pixelgpt Balinese PixelGPT Dataset This dataset contains preprocessed Balinese text data for training PixelGPT models. Dataset Statistics Language: Balinese (bali) Total samples: 54,467 Train samples: 54,017 Test samples: 450 Tokenizers [TO BE EDITED] Features text_id: Document identifier chunk_id: Chunk identifier within document pixel_values: Rendered pixel representation of aksara text <tokenizer_name>_token_ids: Token IDs from grapheme-based… See the full description on the dataset page: https://huggingface.co/datasets/Exqrch/balinese-Komodo-pixelgpt.tabulartext-generation10K<n<100K0 likes75 downloads7mo agoHugging Faceizzako /sundanese-pixelgpt Sundanese PixelGPT Dataset This dataset contains preprocessed Sundanese text data for training PixelGPT models. Dataset Statistics Language: Sundanese (sunda) Total samples: 294,756 Train samples: 293,933 Test samples: 823 Tokenizers Grapheme tokenizer: izzako/sunda-llama-tokenizer LLaMA tokenizer: ernie-research/DualGPT Features text_id: Document identifier chunk_id: Chunk identifier within document pixel_values: Rendered pixel representation… See the full description on the dataset page: https://huggingface.co/datasets/izzako/sundanese-pixelgpt.tabulartext-generation100K<n<1M0 likes72 downloads9mo agoHugging Faceizzako /balinese-pixelgpt Balinese PixelGPT Dataset This dataset contains preprocessed Balinese text data for training PixelGPT models. Dataset Statistics Language: Balinese (bali) Total samples: 54,467 Train samples: 54,017 Test samples: 450 Tokenizers Grapheme tokenizer: izzako/javanese-llama-tokenizer LLaMA tokenizer: ernie-research/DualGPT Features text_id: Document identifier chunk_id: Chunk identifier within document pixel_values: Rendered pixel… See the full description on the dataset page: https://huggingface.co/datasets/izzako/balinese-pixelgpt.tabulartext-generation10K<n<100K0 likes63 downloads9mo agoHugging Face