CoolFace
20 results

macedonian

LVSTCK /macedonian-llm-eval Macedonian LLM Eval This repository is adapted from the original work by Aleksa Gordić. If you find this work useful, please consider citing or acknowledging the original source. You can find the Macedonian LLM eval on GitHub. To run evaluation just follow the guide. Info: You can run the evaluation for Serbian and Slovenian as well, just swap Macedonian with either one of them. What is currently covered: Common sense reasoning: Hellaswag, Winogrande, PIQA… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-llm-eval.1 likes149 downloads1y agoHugging Facemteb /MacedonianTweetSentimentClassification MacedonianTweetSentimentClassification An MTEB dataset Massive Text Embedding Benchmark An Macedonian dataset for tweet sentiment classification. Task category t2c Domains Social, Written Reference https://aclanthology.org/R15-1034/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["MacedonianTweetSentimentClassification"]) evaluator = mteb.MTEB(task) model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/MacedonianTweetSentimentClassification.texttext-classification1K<n<10K0 likes134 downloads1y agoHugging FaceLVSTCK /macedonian-corpus-cleaned-dedup Macedonian Corpus - Cleaned and Deduplicated Paper 🌟 Key Highlights Size: 16.78 GB, Word Count: 1.47 billion Deduplicated using MinHash to remove redundant documents. 📋 Overview Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know, no consolidated resource encompassing all available public data exists. Another challenge is the state of… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-cleaned-dedup.texttext-generation1M<n<10M1 likes56 downloads1y agoHugging FaceLVSTCK /macedonian-corpus-raw Macedonian Corpus - Raw 🌟 Key Highlights Size: 37.6 GB, Word Count: 3.53 billion Includes data from 10+ sources, including academic texts, public archives, and online resources. Minimal preprocessing applied. Examples include academic papers, books, scraped web content, and more. 📋 Overview Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-raw.text10M<n<100M0 likes46 downloads1y agoHugging FaceLVSTCK /macedonian-corpus-cleaned Macedonian Corpus - Cleaned raw version here Paper 🌟 Key Highlights Size: 35.5 GB, Word Count: 3.31 billion Filtered for irrelevant and low-quality content using C4 and Gopher filtering. Includes text from 10+ sources such as fineweb-2, HPLT-2, Wikipedia, and more. 📋 Overview Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know, no consolidated… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-cleaned.texttext-generation1M<n<10M0 likes45 downloads1y agoHugging Faceshunyalabs /macedonian-speech-datasetaudio1K<n<10K0 likes44 downloads1y agoHugging Face