lvs
Datasets
All datasets matching “lvs”macedonian-llm-eval
Macedonian LLM Eval
This repository is adapted from the original work by Aleksa Gordić. If you find this work useful, please consider citing or acknowledging the original source.
You can find the Macedonian LLM eval on GitHub. To run evaluation just follow the guide.
Info: You can run the evaluation for Serbian and Slovenian as well, just swap Macedonian with either one of them.
What is currently covered:
Common sense reasoning: Hellaswag, Winogrande, PIQA… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-llm-eval.lvs6dnq_open-mk
NQ-Open MK version
This dataset is a Macedonian adaptation of the NQ-Open dataset, originally curated (English -> Serbian) by Aleksa Gordić. It was translated from Serbian to Macedonian using the Google Translate API.
You can find this dataset as part of the macedonian-llm-eval GitHub and HuggingFace.
Project page: https://macedonian-llm.github.io/
NOTE: train version of the dataset is not fully complete, as there are about 66k instances instead of 87k (Google Translation API budget… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/nq_open-mk.macedonian-corpus-cleaned-dedup
Macedonian Corpus - Cleaned and Deduplicated
Paper
🌟 Key Highlights
Size: 16.78 GB, Word Count: 1.47 billion
Deduplicated using MinHash to remove redundant documents.
📋 Overview
Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know, no consolidated resource encompassing all available public data exists. Another challenge is the state of… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-cleaned-dedup.LVSM_finalmacedonian-corpus-raw
Macedonian Corpus - Raw
🌟 Key Highlights
Size: 37.6 GB, Word Count: 3.53 billion
Includes data from 10+ sources, including academic texts, public archives, and online resources.
Minimal preprocessing applied.
Examples include academic papers, books, scraped web content, and more.
📋 Overview
Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-raw.
