CoolFace
20 results

Javanese

thesimonharms /public-javanese-dataset Public-Domain Javanese Manuscript & Text Dataset A curated collection of public-domain (or openly-licensed) Javanese-language source material — manuscript scans, plain-text transcriptions, digitized printed books, and aksara Jawa (Carakan) primers. What's in it # Directory Title Material Author / Credit License 1 kakawin-nagarakertagama Kakawin Nagarakertagama (Desawarnnana) Old Javanese (Kawi) kakawin Mpu Prapanca (1365) Public domain 2… See the full description on the dataset page: https://huggingface.co/datasets/thesimonharms/public-javanese-dataset.text-generation100K<n<1M0 likes248 downloads2mo agoHugging Facectaguchi /SLR35_javaneseaudio100K<n<1M0 likes106 downloads23d agoHugging Facew11wo /imdb-javaneseLarge Movie Review Dataset translated to Javanese. This is a dataset for binary sentiment classification containing substantially more data than previous benchmark datasets. We provide a set of 25,000 highly polar movie reviews for training, and 25,000 for testing. There is additional unlabeled data for use as well. We translated the original IMDB Dataset to Javanese using the multi-lingual MarianMT Transformer model from `Helsinki-NLP/opus-mt-en-mul`.texttext-classification100K<n<1M0 likes100 downloads4y agoHugging FaceExqrch /javanese-Komodo-pixelgpt Javanese PixelGPT Dataset This dataset contains preprocessed Javanese text data for training PixelGPT models. Dataset Statistics Language: Javanese (jawa) Total samples: 401,542 Train samples: 400,726 Test samples: 816 Tokenizers [TO BE EDITED] Features text_id: Document identifier chunk_id: Chunk identifier within document pixel_values: Rendered pixel representation of aksara text <tokenizer_name>_token_ids: Token IDs from grapheme-based… See the full description on the dataset page: https://huggingface.co/datasets/Exqrch/javanese-Komodo-pixelgpt.tabulartext-generation100K<n<1M0 likes100 downloads7mo agoHugging FaceExqrch /Rebuttal-javanese-pixelgpt Javanese PixelGPT Tokenizer Ablation Dataset Optimized with Font Size 6 and Dynamic Trimming. Tokenizer Schema tok_grapheme: Language-specific Grapheme BPE (izzako/javanese-llama-tokenizer) tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT) tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1) tok_mt5: Google Multilingual Unigram (google/mt5-small) tabular100K<n<1M0 likes83 downloads6mo agoHugging Facesaillab /alpaca-javanese-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-javanese-cleaned.text10K<n<100K0 likes79 downloads2y agoHugging Face