CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01thesimonharms /public-javanese-dataset Public-Domain Javanese Manuscript & Text Dataset A curated collection of public-domain (or openly-licensed) Javanese-language source material — manuscript scans, plain-text transcriptions, digitized printed books, and aksara Jawa (Carakan) primers. What's in it # Directory Title Material Author / Credit License 1 kakawin-nagarakertagama Kakawin Nagarakertagama (Desawarnnana) Old Javanese (Kawi) kakawin Mpu Prapanca (1365) Public domain 2… See the full description on the dataset page: https://huggingface.co/datasets/thesimonharms/public-javanese-dataset.text-generation100K<n<1M0 likes248 downloads2mo agoHugging Face02ctaguchi /SLR35_javaneseaudio100K<n<1M0 likes106 downloads23d agoHugging Face03w11wo /imdb-javaneseLarge Movie Review Dataset translated to Javanese. This is a dataset for binary sentiment classification containing substantially more data than previous benchmark datasets. We provide a set of 25,000 highly polar movie reviews for training, and 25,000 for testing. There is additional unlabeled data for use as well. We translated the original IMDB Dataset to Javanese using the multi-lingual MarianMT Transformer model from `Helsinki-NLP/opus-mt-en-mul`.texttext-classification100K<n<1M0 likes100 downloads4y agoHugging Face04Exqrch /javanese-Komodo-pixelgpt Javanese PixelGPT Dataset This dataset contains preprocessed Javanese text data for training PixelGPT models. Dataset Statistics Language: Javanese (jawa) Total samples: 401,542 Train samples: 400,726 Test samples: 816 Tokenizers [TO BE EDITED] Features text_id: Document identifier chunk_id: Chunk identifier within document pixel_values: Rendered pixel representation of aksara text <tokenizer_name>_token_ids: Token IDs from grapheme-based… See the full description on the dataset page: https://huggingface.co/datasets/Exqrch/javanese-Komodo-pixelgpt.tabulartext-generation100K<n<1M0 likes100 downloads7mo agoHugging Face05Exqrch /Rebuttal-javanese-pixelgpt Javanese PixelGPT Tokenizer Ablation Dataset Optimized with Font Size 6 and Dynamic Trimming. Tokenizer Schema tok_grapheme: Language-specific Grapheme BPE (izzako/javanese-llama-tokenizer) tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT) tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1) tok_mt5: Google Multilingual Unigram (google/mt5-small) tabular100K<n<1M0 likes83 downloads6mo agoHugging Face06saillab /alpaca-javanese-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-javanese-cleaned.text10K<n<100K0 likes79 downloads2y agoHugging Face07afrizalha /Gatra-1-Javanese GatraOne (Gatra-1) is a synthethic Jawa Krama instruction-tuning dataset, generated by GPT-4. Introducing the Gatra-1 dataset This is a synthetic dataset to fine-tune LLMs into responding in Jawa Krama, the high-register of Javanese language. It is 98% generated using GPT-4, which has very good Jawa Krama capabilities. It is currently a 'beta' version with only 560 input-output prompts. So far, this has been only tested on fine-tuning GPT-3.5 with considerable success.… See the full description on the dataset page: https://huggingface.co/datasets/afrizalha/Gatra-1-Javanese.texttext-generationn<1K3 likes74 downloads2y agoHugging Face08JavaneseHonorifics /Unggah-Ungguh Javanese Honorifics Dataset (Unggah-Ungguh - Released Version) The Javanese language, spoken by over 98 million people, features a distinctive honorific system known as Unggah-Ungguh Basa. In this dataset we present UNGGAH-UNGGUH, a carefully curated dataset designed to encapsulate the nuances of Unggah-Ungguh Basa, the Javanese speech etiquette framework that dictates the choice of words and phrases based on social hierarchy and context. Paper: https://arxiv.org/pdf/2502.20864… See the full description on the dataset page: https://huggingface.co/datasets/JavaneseHonorifics/Unggah-Ungguh.tabular1K<n<10K0 likes38 downloads7mo agoHugging Face09rifoag /javanese_sundanese_story_clozetext1K<n<10K0 likes36 downloads2y agoHugging Face10shunyalabs /javanese-speech-datasetaudio1K<n<10K0 likes34 downloads1y agoHugging Face11cheikh1499 /javaneseaudio1K<n<10K0 likes31 downloads5d agoHugging Face12DGurgurov /javanese_sa Sentiment Analysis Data for the Javanese Language Dataset Description: This dataset contains a sentiment analysis data from Wongso et al. (2021). Data Structure: The data was used for the project on injecting external commonsense knowledge into multilingual Large Language Models. Citation: @inproceedings{wongso2021causal, title={Causal and Masked Language Modeling of Javanese Language using Transformer-based Architectures}, author={Wongso, Wilson and Setiawan, David Samuel… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/javanese_sa.texttext-classification10K<n<100K0 likes28 downloads2y agoHugging Face13saillab /alpaca_javanese_tacoThis repository contains the dataset used for the TaCo paper. The dataset follows the style outlined in the TaCo paper, as follows: { "instruction": "instruction in xx", "input": "input in xx", "output": "Instruction in English: instruction in en , Response in English: response in en , Response in xx: response in xx " } Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_javanese_taco.text10K<n<100K0 likes27 downloads2y agoHugging Face14DGurgurov /javanese_conceptnet ConceptNet Data for the Javanese Language Dataset Description: This dataset contains data extracted from ConceptNet using the dedicated module for fetching knowledge from the graph, available on GitHub. Data Structure: The data is converted from triplets into natural text using a pre-defined relationship mapping and split into training and validation sets. It was used for training language adapters for the project aimed at injecting external commonsense knowledge into multilingual… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/javanese_conceptnet.text1K<n<10K1 likes26 downloads2y agoHugging Face15itsqy /javanese-to-indonesia0 likes26 downloads1mo agoHugging Face16mteb /JavaneseIMDBClassification JavaneseIMDBClassification An MTEB dataset Massive Text Embedding Benchmark Large Movie Review Dataset translated to Javanese. This is a dataset for binary sentiment classification containing substantially more data than previous benchmark datasets. Task category t2c Domains Reviews, Written Reference https://github.com/w11wo/nlp-datasets#javanese-imdb How to evaluate on this task You can evaluate an embedding model on this dataset using the following… See the full description on the dataset page: https://huggingface.co/datasets/mteb/JavaneseIMDBClassification.texttext-classification10K<n<100K0 likes25 downloads1y agoHugging Face17afrizalha /Centhini-1-Javanese Dataset details The dataset comprises 529,575 pretraining examples for both Ngoko and Krama Javanese. The data is almost predominantly translation generated with Deepseek V3. The translations include data from English language Fineweb and a paraphrased translation from Indonesian mc4 dataset. Other examples here include ancient Javanese texts, like Serat Centhini and Babad Tanah Djawi, but also open texts like Javanese wikipedia. To our knowledge, this is the largest easily… See the full description on the dataset page: https://huggingface.co/datasets/afrizalha/Centhini-1-Javanese.texttext-generation100K<n<1M3 likes24 downloads2y agoHugging Face18ravialdy /javanese-translatedtabular100K<n<1M0 likes21 downloads3y agoHugging Face19afrizalha /Gatra-2-Javanese Dataset details The dataset comprises 36870 prompt-response pairs of Krama Javanese instruction-tuning examples. The data is almost entirely synthetic with minimal human curation. The current dataset supports only single-turn QA, although fine-tuning on instruction-tuned models may allow for transfer of multi-turn capabilities. The prompts are generated by GPT-4o, while the responses are generated by Claude 3 Haiku. The way the data set was generated, the prompt may contain terms in… See the full description on the dataset page: https://huggingface.co/datasets/afrizalha/Gatra-2-Javanese.texttext-generation10K<n<100K3 likes21 downloads2y agoHugging Face20rmdhirr /alpaca-javanese-7k-cleanedtext1K<n<10K1 likes21 downloads1y agoHugging Face21Speech-data /Javanese-Speech-Dataset 🎧 Javanese Speech Dataset The Javanese Speech Dataset is a structured and scalable speech audio dataset designed to provide high-quality audio data for training modern AI and machine learning models. It includes 85 hours of audio data across 585 files, delivered in MP3 and WAV formats, with a total size of 104 MB. This well-balanced audio dataset offers diverse and representative voice data, with 51% female and 49% male speakers, and an age range spanning from 18 to 50+ years. The… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Javanese-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes21 downloads6mo agoHugging Face22richardcsuwandi /oasst-javanese Dataset Summary We translated the OpenAssistant Conversations (OASST) dataset into Javanese using Meta's No Language Left Behind (NLLB) model. Why Javanese? Javanese is spoken by over 90 million people on the island of Java in Indonesia. While its prevalence is comparable to other widely spoken languages, such as Vietnamese and Turkish, its representation in current large language model (LLM) chatbots remains limited. By translating this dataset, we aim to enhance the… See the full description on the dataset page: https://huggingface.co/datasets/richardcsuwandi/oasst-javanese.tabularquestion-answering100K<n<1M0 likes17 downloads2y agoHugging Face23octava /fork-google-openslr-javaneseaudio1K<n<10K0 likes17 downloads2y agoHugging Face24junwatu /javanese-multilingual-lexicongated Javanese Multilingual Lexicon A multilingual lexicon dataset of 202,912 entries derived from Sastra.org, a digital archive of Javanese literary heritage maintained by Yayasan Sastra Lestari. The dataset compiles 32 historical lexicographic sources spanning from the 1830s to 2010, covering Javanese, Old Javanese (Kawi), Indonesian, Dutch, English, and French. Dataset Structure The dataset is provided as JSONL files. The full corpus is in data/leksikon_all.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/junwatu/javanese-multilingual-lexicon.texttranslation100K<n<1M1 likes17 downloads2mo agoHugging Face25InfoBayAI /Javanese-Non-STEM-Textbook-DatasetgatedDataset Description: This dataset is a large-scale collection of Javanese Non-STEM textbook data, containing 597 books and 67.05 million words, designed to support the development and training of advanced NLP systems and AI models for language understanding, reasoning, and general knowledge learning in Javanese. Full Dataset Overview This dataset is part of a large-scale multilingual educational corpus containing over 3+ billion words across 5,000+ subjects, supported by interwoven images for… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Javanese-Non-STEM-Textbook-Dataset.tabular10K<n<100K0 likes17 downloads6d agoHugging Face26mesolitica /ms-javanese0 likes16 downloads4y agoHugging Face27biznetgio /alpaca-clean-javanesetext10K<n<100K2 likes14 downloads2y agoHugging Face28izzako /javanese-pixelgpt Javanese PixelGPT Dataset This dataset contains preprocessed Javanese text data for training PixelGPT models. Dataset Statistics Language: Javanese (jawa) Total samples: 401,542 Train samples: 400,726 Test samples: 816 Tokenizers Grapheme tokenizer: izzako/javanese-llama-tokenizer LLaMA tokenizer: ernie-research/DualGPT Features text_id: Document identifier chunk_id: Chunk identifier within document pixel_values: Rendered pixel… See the full description on the dataset page: https://huggingface.co/datasets/izzako/javanese-pixelgpt.tabulartext-generation100K<n<1M0 likes12 downloads9mo agoHugging Face29afrizalha /javanese-collectiontext100K<n<1M0 likes10 downloads2y agoHugging Face30irasalsabila /javanese_asr_dataset_20ktext10K<n<100K0 likes10 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.