datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
javanese-Komodo-pixelgpt
Javanese PixelGPT Dataset
This dataset contains preprocessed Javanese text data for training PixelGPT models.
Dataset Statistics
Language: Javanese (jawa)
Total samples: 401,542
Train samples: 400,726
Test samples: 816
Tokenizers
[TO BE EDITED]
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel representation of aksara text
<tokenizer_name>_token_ids: Token IDs from grapheme-based… See the full description on the dataset page: https://huggingface.co/datasets/Exqrch/javanese-Komodo-pixelgpt.Rebuttal-javanese-pixelgpt
Javanese PixelGPT Tokenizer Ablation Dataset
Optimized with Font Size 6 and Dynamic Trimming.
Tokenizer Schema
tok_grapheme: Language-specific Grapheme BPE (izzako/javanese-llama-tokenizer)
tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT)
tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1)
tok_mt5: Google Multilingual Unigram (google/mt5-small)
balinese-Komodo-pixelgpt
Balinese PixelGPT Dataset
This dataset contains preprocessed Balinese text data for training PixelGPT models.
Dataset Statistics
Language: Balinese (bali)
Total samples: 54,467
Train samples: 54,017
Test samples: 450
Tokenizers
[TO BE EDITED]
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel representation of aksara text
<tokenizer_name>_token_ids: Token IDs from grapheme-based… See the full description on the dataset page: https://huggingface.co/datasets/Exqrch/balinese-Komodo-pixelgpt.sundanese-pixelgpt
Sundanese PixelGPT Dataset
This dataset contains preprocessed Sundanese text data for training PixelGPT models.
Dataset Statistics
Language: Sundanese (sunda)
Total samples: 294,756
Train samples: 293,933
Test samples: 823
Tokenizers
Grapheme tokenizer: izzako/sunda-llama-tokenizer
LLaMA tokenizer: ernie-research/DualGPT
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel representation… See the full description on the dataset page: https://huggingface.co/datasets/izzako/sundanese-pixelgpt.balinese-pixelgpt
Balinese PixelGPT Dataset
This dataset contains preprocessed Balinese text data for training PixelGPT models.
Dataset Statistics
Language: Balinese (bali)
Total samples: 54,467
Train samples: 54,017
Test samples: 450
Tokenizers
Grapheme tokenizer: izzako/javanese-llama-tokenizer
LLaMA tokenizer: ernie-research/DualGPT
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel… See the full description on the dataset page: https://huggingface.co/datasets/izzako/balinese-pixelgpt.Rebuttal-sundanese-pixelgpt
Sundanese PixelGPT Tokenizer Ablation Dataset
Optimized with Font Size 6 and Dynamic Trimming.
Tokenizer Schema
tok_grapheme: Language-specific Grapheme BPE (izzako/sunda-llama-tokenizer)
tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT)
tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1)
tok_mt5: Google Multilingual Unigram (google/mt5-small)
lampung-pixelgpt
Lampung PixelGPT Dataset
This dataset contains preprocessed Lampung text data for training PixelGPT models.
Dataset Statistics
Language: Lampung (lampung)
Total samples: 1,029
Train samples: 945
Test samples: 84
Tokenizers
Grapheme tokenizer: izzako/sunda-llama-tokenizer
LLaMA tokenizer: ernie-research/DualGPT
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel representation of aksara… See the full description on the dataset page: https://huggingface.co/datasets/izzako/lampung-pixelgpt.javanese-pixelgpt
Javanese PixelGPT Dataset
This dataset contains preprocessed Javanese text data for training PixelGPT models.
Dataset Statistics
Language: Javanese (jawa)
Total samples: 401,542
Train samples: 400,726
Test samples: 816
Tokenizers
Grapheme tokenizer: izzako/javanese-llama-tokenizer
LLaMA tokenizer: ernie-research/DualGPT
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel… See the full description on the dataset page: https://huggingface.co/datasets/izzako/javanese-pixelgpt.Rebuttal-sundanese-pixelgpt-debug
Sundanese PixelGPT Tokenizer Ablation Dataset
Optimized with Font Size 6 and Dynamic Trimming.
Tokenizer Schema
tok_grapheme: Language-specific Grapheme BPE (izzako/sunda-llama-tokenizer)
tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT)
tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1)
tok_mt5: Google Multilingual Unigram (google/mt5-small)
Rebuttal-lampung-pixelgpt
Lampung PixelGPT Tokenizer Ablation Dataset
Optimized with Font Size 6 and Dynamic Trimming.
Tokenizer Schema
tok_grapheme: Language-specific Grapheme BPE (izzako/sunda-llama-tokenizer)
tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT)
tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1)
tok_mt5: Google Multilingual Unigram (google/mt5-small)
Rebuttal-javanese-pixelgpt-debug
Javanese PixelGPT Tokenizer Ablation Dataset
Optimized with Font Size 6 and Dynamic Trimming.
Tokenizer Schema
tok_grapheme: Language-specific Grapheme BPE (izzako/javanese-llama-tokenizer)
tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT)
tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1)
tok_mt5: Google Multilingual Unigram (google/mt5-small)
Rebuttal-lampung-pixelgpt-debug
Lampung PixelGPT Tokenizer Ablation Dataset
Optimized with Font Size 6 and Dynamic Trimming.
Tokenizer Schema
tok_grapheme: Language-specific Grapheme BPE (izzako/sunda-llama-tokenizer)
tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT)
tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1)
tok_mt5: Google Multilingual Unigram (google/mt5-small)
Rebuttal-balinese-pixelgpt
Balinese PixelGPT Tokenizer Ablation Dataset
Optimized with Font Size 6 and Dynamic Trimming.
Tokenizer Schema
tok_grapheme: Language-specific Grapheme BPE (izzako/javanese-llama-tokenizer)
tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT)
tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1)
tok_mt5: Google Multilingual Unigram (google/mt5-small)
Rebuttal-balinese-pixelgpt-debug
Balinese PixelGPT Tokenizer Ablation Dataset
Optimized with Font Size 6 and Dynamic Trimming.
Tokenizer Schema
tok_grapheme: Language-specific Grapheme BPE (izzako/javanese-llama-tokenizer)
tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT)
tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1)
tok_mt5: Google Multilingual Unigram (google/mt5-small)
