CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Exqrch /javanese-Komodo-pixelgpt Javanese PixelGPT Dataset This dataset contains preprocessed Javanese text data for training PixelGPT models. Dataset Statistics Language: Javanese (jawa) Total samples: 401,542 Train samples: 400,726 Test samples: 816 Tokenizers [TO BE EDITED] Features text_id: Document identifier chunk_id: Chunk identifier within document pixel_values: Rendered pixel representation of aksara text <tokenizer_name>_token_ids: Token IDs from grapheme-based… See the full description on the dataset page: https://huggingface.co/datasets/Exqrch/javanese-Komodo-pixelgpt.tabulartext-generation100K<n<1M0 likes96 downloads7mo agoHugging Face02afrizalha /Gatra-1-Javanese GatraOne (Gatra-1) is a synthethic Jawa Krama instruction-tuning dataset, generated by GPT-4. Introducing the Gatra-1 dataset This is a synthetic dataset to fine-tune LLMs into responding in Jawa Krama, the high-register of Javanese language. It is 98% generated using GPT-4, which has very good Jawa Krama capabilities. It is currently a 'beta' version with only 560 input-output prompts. So far, this has been only tested on fine-tuning GPT-3.5 with considerable success.… See the full description on the dataset page: https://huggingface.co/datasets/afrizalha/Gatra-1-Javanese.texttext-generationn<1K3 likes79 downloads2y agoHugging Face03afrizalha /Centhini-1-Javanese Dataset details The dataset comprises 529,575 pretraining examples for both Ngoko and Krama Javanese. The data is almost predominantly translation generated with Deepseek V3. The translations include data from English language Fineweb and a paraphrased translation from Indonesian mc4 dataset. Other examples here include ancient Javanese texts, like Serat Centhini and Babad Tanah Djawi, but also open texts like Javanese wikipedia. To our knowledge, this is the largest easily… See the full description on the dataset page: https://huggingface.co/datasets/afrizalha/Centhini-1-Javanese.texttext-generation100K<n<1M3 likes27 downloads2y agoHugging Face04afrizalha /Gatra-2-Javanese Dataset details The dataset comprises 36870 prompt-response pairs of Krama Javanese instruction-tuning examples. The data is almost entirely synthetic with minimal human curation. The current dataset supports only single-turn QA, although fine-tuning on instruction-tuned models may allow for transfer of multi-turn capabilities. The prompts are generated by GPT-4o, while the responses are generated by Claude 3 Haiku. The way the data set was generated, the prompt may contain terms in… See the full description on the dataset page: https://huggingface.co/datasets/afrizalha/Gatra-2-Javanese.texttext-generation10K<n<100K3 likes26 downloads2y agoHugging Face05junwatu /javanese-multilingual-lexicongated Javanese Multilingual Lexicon A multilingual lexicon dataset of 202,912 entries derived from Sastra.org, a digital archive of Javanese literary heritage maintained by Yayasan Sastra Lestari. The dataset compiles 32 historical lexicographic sources spanning from the 1830s to 2010, covering Javanese, Old Javanese (Kawi), Indonesian, Dutch, English, and French. Dataset Structure The dataset is provided as JSONL files. The full corpus is in data/leksikon_all.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/junwatu/javanese-multilingual-lexicon.texttranslation100K<n<1M1 likes16 downloads2mo agoHugging Face06izzako /javanese-pixelgpt Javanese PixelGPT Dataset This dataset contains preprocessed Javanese text data for training PixelGPT models. Dataset Statistics Language: Javanese (jawa) Total samples: 401,542 Train samples: 400,726 Test samples: 816 Tokenizers Grapheme tokenizer: izzako/javanese-llama-tokenizer LLaMA tokenizer: ernie-research/DualGPT Features text_id: Document identifier chunk_id: Chunk identifier within document pixel_values: Rendered pixel… See the full description on the dataset page: https://huggingface.co/datasets/izzako/javanese-pixelgpt.tabulartext-generation100K<n<1M0 likes11 downloads9mo agoHugging Face07Iftitahu /javanese_instruct_storiesgatedA dataset of parallel translation-based instructions for Javanese language as a target language. Materials are taken from randomly selected children stories at https://storyweaver.org.in, under CC-By-SA-4.0 license. The template IDs are: (1, 'Terjemahno penggalan teks crito ing ngisor iki saka Bahasa Inggris dadi teks crito ing Basa Jawa:', 'Terjemahane utawa padanan teks crito kasebut ing Basa Jawa yaiku:'), (2, 'Terjemahno penggalan teks crito ing ngisor iki saka Bahasa Indonesia dadi teks… See the full description on the dataset page: https://huggingface.co/datasets/Iftitahu/javanese_instruct_stories.texttranslationn<1K2 likes6 downloads3y agoHugging Face08rinnieyoung /sea-javanese-cleaned-parquet-v1 SEA Javanese Cleaned Parquet v1 Dataset Summary This dataset is a cleaned Javanese pretraining corpus exported in Hugging Face parquet format. Current public sources used in this release: HuggingFaceFW/fineweb-2 / jav_Latn allenai/c4 / jv afrizalha/Centhini-1-Javanese Cleaning and Deduplication Current pipeline: basic text cleaning short-text filtering repetition filtering rule-based noise filtering document-level exact deduplication across all included… See the full description on the dataset page: https://huggingface.co/datasets/rinnieyoung/sea-javanese-cleaned-parquet-v1.texttext-generation100K<n<1M0 likes6 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.