macedonian
macedonian-llm-eval
Macedonian LLM Eval
This repository is adapted from the original work by Aleksa Gordić. If you find this work useful, please consider citing or acknowledging the original source.
You can find the Macedonian LLM eval on GitHub. To run evaluation just follow the guide.
Info: You can run the evaluation for Serbian and Slovenian as well, just swap Macedonian with either one of them.
What is currently covered:
Common sense reasoning: Hellaswag, Winogrande, PIQA… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-llm-eval.MacedonianTweetSentimentClassification
MacedonianTweetSentimentClassification
An MTEB dataset
Massive Text Embedding Benchmark
An Macedonian dataset for tweet sentiment classification.
Task category
t2c
Domains
Social, Written
Reference
https://aclanthology.org/R15-1034/
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["MacedonianTweetSentimentClassification"])
evaluator = mteb.MTEB(task)
model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/MacedonianTweetSentimentClassification.macedonian-corpus-cleaned-dedup
Macedonian Corpus - Cleaned and Deduplicated
Paper
🌟 Key Highlights
Size: 16.78 GB, Word Count: 1.47 billion
Deduplicated using MinHash to remove redundant documents.
📋 Overview
Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know, no consolidated resource encompassing all available public data exists. Another challenge is the state of… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-cleaned-dedup.macedonian-corpus-raw
Macedonian Corpus - Raw
🌟 Key Highlights
Size: 37.6 GB, Word Count: 3.53 billion
Includes data from 10+ sources, including academic texts, public archives, and online resources.
Minimal preprocessing applied.
Examples include academic papers, books, scraped web content, and more.
📋 Overview
Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-raw.macedonian-corpus-cleaned
Macedonian Corpus - Cleaned
raw version here
Paper
🌟 Key Highlights
Size: 35.5 GB, Word Count: 3.31 billion
Filtered for irrelevant and low-quality content using C4 and Gopher filtering.
Includes text from 10+ sources such as fineweb-2, HPLT-2, Wikipedia, and more.
📋 Overview
Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know, no consolidated… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-cleaned.macedonian-speech-dataset
