CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01unstonio /pixelgpt-24x24-20k PixelGPT 24×24 — 20K 20,000 native 24×24 pixel-art sprites with captions and semantic taxonomy labels. This is a clean, rebalanced, rights-conscious public subset of the larger PixelGPT 24×24 dataset. Every sprite: is rendered at a native resolution of 24×24 pixels uses no more than 5 colors includes an original text caption is assigned to a two-level semantic taxonomy is distributed in lossless PNG and Parquet formats Looking for the complete dataset?The full edition… See the full description on the dataset page: https://huggingface.co/datasets/unstonio/pixelgpt-24x24-20k.imagetext-to-image10K<n<100K69 likes539 downloads2mo agoHugging Face02Exqrch /javanese-Komodo-pixelgpt Javanese PixelGPT Dataset This dataset contains preprocessed Javanese text data for training PixelGPT models. Dataset Statistics Language: Javanese (jawa) Total samples: 401,542 Train samples: 400,726 Test samples: 816 Tokenizers [TO BE EDITED] Features text_id: Document identifier chunk_id: Chunk identifier within document pixel_values: Rendered pixel representation of aksara text <tokenizer_name>_token_ids: Token IDs from grapheme-based… See the full description on the dataset page: https://huggingface.co/datasets/Exqrch/javanese-Komodo-pixelgpt.tabulartext-generation100K<n<1M0 likes98 downloads7mo agoHugging Face03Exqrch /Rebuttal-javanese-pixelgpt Javanese PixelGPT Tokenizer Ablation Dataset Optimized with Font Size 6 and Dynamic Trimming. Tokenizer Schema tok_grapheme: Language-specific Grapheme BPE (izzako/javanese-llama-tokenizer) tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT) tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1) tok_mt5: Google Multilingual Unigram (google/mt5-small) tabular100K<n<1M0 likes89 downloads6mo agoHugging Face04Exqrch /balinese-Komodo-pixelgpt Balinese PixelGPT Dataset This dataset contains preprocessed Balinese text data for training PixelGPT models. Dataset Statistics Language: Balinese (bali) Total samples: 54,467 Train samples: 54,017 Test samples: 450 Tokenizers [TO BE EDITED] Features text_id: Document identifier chunk_id: Chunk identifier within document pixel_values: Rendered pixel representation of aksara text <tokenizer_name>_token_ids: Token IDs from grapheme-based… See the full description on the dataset page: https://huggingface.co/datasets/Exqrch/balinese-Komodo-pixelgpt.tabulartext-generation10K<n<100K0 likes74 downloads7mo agoHugging Face05izzako /sundanese-pixelgpt Sundanese PixelGPT Dataset This dataset contains preprocessed Sundanese text data for training PixelGPT models. Dataset Statistics Language: Sundanese (sunda) Total samples: 294,756 Train samples: 293,933 Test samples: 823 Tokenizers Grapheme tokenizer: izzako/sunda-llama-tokenizer LLaMA tokenizer: ernie-research/DualGPT Features text_id: Document identifier chunk_id: Chunk identifier within document pixel_values: Rendered pixel representation… See the full description on the dataset page: https://huggingface.co/datasets/izzako/sundanese-pixelgpt.tabulartext-generation100K<n<1M0 likes70 downloads9mo agoHugging Face06izzako /balinese-pixelgpt Balinese PixelGPT Dataset This dataset contains preprocessed Balinese text data for training PixelGPT models. Dataset Statistics Language: Balinese (bali) Total samples: 54,467 Train samples: 54,017 Test samples: 450 Tokenizers Grapheme tokenizer: izzako/javanese-llama-tokenizer LLaMA tokenizer: ernie-research/DualGPT Features text_id: Document identifier chunk_id: Chunk identifier within document pixel_values: Rendered pixel… See the full description on the dataset page: https://huggingface.co/datasets/izzako/balinese-pixelgpt.tabulartext-generation10K<n<100K0 likes62 downloads9mo agoHugging Face07Exqrch /Rebuttal-sundanese-pixelgpt Sundanese PixelGPT Tokenizer Ablation Dataset Optimized with Font Size 6 and Dynamic Trimming. Tokenizer Schema tok_grapheme: Language-specific Grapheme BPE (izzako/sunda-llama-tokenizer) tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT) tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1) tok_mt5: Google Multilingual Unigram (google/mt5-small) tabular100K<n<1M0 likes47 downloads6mo agoHugging Face08aldigobbler /chatml-pixelgpt-24x24-20ktext10K<n<100K0 likes26 downloads2mo agoHugging Face09izzako /lampung-pixelgpt Lampung PixelGPT Dataset This dataset contains preprocessed Lampung text data for training PixelGPT models. Dataset Statistics Language: Lampung (lampung) Total samples: 1,029 Train samples: 945 Test samples: 84 Tokenizers Grapheme tokenizer: izzako/sunda-llama-tokenizer LLaMA tokenizer: ernie-research/DualGPT Features text_id: Document identifier chunk_id: Chunk identifier within document pixel_values: Rendered pixel representation of aksara… See the full description on the dataset page: https://huggingface.co/datasets/izzako/lampung-pixelgpt.tabulartext-generation1K<n<10K0 likes13 downloads9mo agoHugging Face10izzako /javanese-pixelgpt Javanese PixelGPT Dataset This dataset contains preprocessed Javanese text data for training PixelGPT models. Dataset Statistics Language: Javanese (jawa) Total samples: 401,542 Train samples: 400,726 Test samples: 816 Tokenizers Grapheme tokenizer: izzako/javanese-llama-tokenizer LLaMA tokenizer: ernie-research/DualGPT Features text_id: Document identifier chunk_id: Chunk identifier within document pixel_values: Rendered pixel… See the full description on the dataset page: https://huggingface.co/datasets/izzako/javanese-pixelgpt.tabulartext-generation100K<n<1M0 likes11 downloads9mo agoHugging Face11Exqrch /Rebuttal-sundanese-pixelgpt-debug Sundanese PixelGPT Tokenizer Ablation Dataset Optimized with Font Size 6 and Dynamic Trimming. Tokenizer Schema tok_grapheme: Language-specific Grapheme BPE (izzako/sunda-llama-tokenizer) tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT) tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1) tok_mt5: Google Multilingual Unigram (google/mt5-small) tabular1K<n<10K0 likes11 downloads6mo agoHugging Face12Exqrch /Rebuttal-javanese-pixelgpt-debug Javanese PixelGPT Tokenizer Ablation Dataset Optimized with Font Size 6 and Dynamic Trimming. Tokenizer Schema tok_grapheme: Language-specific Grapheme BPE (izzako/javanese-llama-tokenizer) tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT) tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1) tok_mt5: Google Multilingual Unigram (google/mt5-small) tabular1K<n<10K0 likes8 downloads6mo agoHugging Face13Exqrch /Rebuttal-lampung-pixelgpt Lampung PixelGPT Tokenizer Ablation Dataset Optimized with Font Size 6 and Dynamic Trimming. Tokenizer Schema tok_grapheme: Language-specific Grapheme BPE (izzako/sunda-llama-tokenizer) tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT) tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1) tok_mt5: Google Multilingual Unigram (google/mt5-small) tabular1K<n<10K0 likes8 downloads6mo agoHugging Face14Exqrch /Rebuttal-lampung-pixelgpt-debug Lampung PixelGPT Tokenizer Ablation Dataset Optimized with Font Size 6 and Dynamic Trimming. Tokenizer Schema tok_grapheme: Language-specific Grapheme BPE (izzako/sunda-llama-tokenizer) tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT) tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1) tok_mt5: Google Multilingual Unigram (google/mt5-small) tabular1K<n<10K0 likes7 downloads6mo agoHugging Face15Exqrch /Rebuttal-balinese-pixelgpt Balinese PixelGPT Tokenizer Ablation Dataset Optimized with Font Size 6 and Dynamic Trimming. Tokenizer Schema tok_grapheme: Language-specific Grapheme BPE (izzako/javanese-llama-tokenizer) tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT) tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1) tok_mt5: Google Multilingual Unigram (google/mt5-small) tabular10K<n<100K0 likes5 downloads6mo agoHugging Face16Exqrch /Rebuttal-balinese-pixelgpt-debug Balinese PixelGPT Tokenizer Ablation Dataset Optimized with Font Size 6 and Dynamic Trimming. Tokenizer Schema tok_grapheme: Language-specific Grapheme BPE (izzako/javanese-llama-tokenizer) tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT) tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1) tok_mt5: Google Multilingual Unigram (google/mt5-small) tabular1K<n<10K0 likes3 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.