CoolFace
20 results

lvs

LVSTCK /macedonian-llm-eval Macedonian LLM Eval This repository is adapted from the original work by Aleksa Gordić. If you find this work useful, please consider citing or acknowledging the original source. You can find the Macedonian LLM eval on GitHub. To run evaluation just follow the guide. Info: You can run the evaluation for Serbian and Slovenian as well, just swap Macedonian with either one of them. What is currently covered: Common sense reasoning: Hellaswag, Winogrande, PIQA… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-llm-eval.1 likes143 downloads1y agoHugging Facexingyoujun /lvs6d0 likes74 downloads6mo agoHugging FaceLVSTCK /nq_open-mk NQ-Open MK version This dataset is a Macedonian adaptation of the NQ-Open dataset, originally curated (English -> Serbian) by Aleksa Gordić. It was translated from Serbian to Macedonian using the Google Translate API. You can find this dataset as part of the macedonian-llm-eval GitHub and HuggingFace. Project page: https://macedonian-llm.github.io/ NOTE: train version of the dataset is not fully complete, as there are about 66k instances instead of 87k (Google Translation API budget… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/nq_open-mk.textquestion-answering10K<n<100K0 likes63 downloads1y agoHugging FaceLVSTCK /macedonian-corpus-cleaned-dedup Macedonian Corpus - Cleaned and Deduplicated Paper 🌟 Key Highlights Size: 16.78 GB, Word Count: 1.47 billion Deduplicated using MinHash to remove redundant documents. 📋 Overview Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know, no consolidated resource encompassing all available public data exists. Another challenge is the state of… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-cleaned-dedup.texttext-generation1M<n<10M1 likes55 downloads1y agoHugging FaceLinaSad /LVSM_finaltext100K<n<1M0 likes49 downloads1y agoHugging FaceLVSTCK /macedonian-corpus-raw Macedonian Corpus - Raw 🌟 Key Highlights Size: 37.6 GB, Word Count: 3.53 billion Includes data from 10+ sources, including academic texts, public archives, and online resources. Minimal preprocessing applied. Examples include academic papers, books, scraped web content, and more. 📋 Overview Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-raw.text10M<n<100M0 likes45 downloads1y agoHugging Face