datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MacedonianTweetSentimentClassification
MacedonianTweetSentimentClassification
An MTEB dataset
Massive Text Embedding Benchmark
An Macedonian dataset for tweet sentiment classification.
Task category
t2c
Domains
Social, Written
Reference
https://aclanthology.org/R15-1034/
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["MacedonianTweetSentimentClassification"])
evaluator = mteb.MTEB(task)
model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/MacedonianTweetSentimentClassification.Muon-MACE-data
Muon-MACE: selection inputs, benchmark results and figure data
Data companion to Muon-MACE, a
MACE research implementation with hybrid Muon–Adam optimization.
This repository contains 52 individually accessible data files (220,102,043 bytes).
Browse a table, download an array or clone the directory tree; no ZIP extraction
is needed. Code, configurations and plotting programs live on GitHub.
中文指南 · File catalog ·
Checksums and download mapping · Data dictionary
Find… See the full description on the dataset page: https://huggingface.co/datasets/shiqiao123/Muon-MACE-data.MACE
MACE Dataset
Dataset for paper "Evaluating and Calibrating LLM Confidence on Questions with Multiple Correct Answers".
MACE-Dance
🎵 MACE-Dance Dataset
MACE-Dance is a large-scale dataset for music-driven dance video generation, released with our SIGGRAPH 2026 paper:
MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generation
It is designed to support research on generating dance videos that are both:
🕺 kinematically plausible
🎨 visually coherent
🎼 well aligned with music
✨ Overview
The dataset contains approximately:
70K dance video clips
5–10 seconds per… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/MACE-Dance.macedonian-corpus-cleaned-dedup
Macedonian Corpus - Cleaned and Deduplicated
Paper
🌟 Key Highlights
Size: 16.78 GB, Word Count: 1.47 billion
Deduplicated using MinHash to remove redundant documents.
📋 Overview
Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know, no consolidated resource encompassing all available public data exists. Another challenge is the state of… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-cleaned-dedup.macedonian-corpus-raw
Macedonian Corpus - Raw
🌟 Key Highlights
Size: 37.6 GB, Word Count: 3.53 billion
Includes data from 10+ sources, including academic texts, public archives, and online resources.
Minimal preprocessing applied.
Examples include academic papers, books, scraped web content, and more.
📋 Overview
Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-raw.macedonian-corpus-cleaned
Macedonian Corpus - Cleaned
raw version here
Paper
🌟 Key Highlights
Size: 35.5 GB, Word Count: 3.31 billion
Filtered for irrelevant and low-quality content using C4 and Gopher filtering.
Includes text from 10+ sources such as fineweb-2, HPLT-2, Wikipedia, and more.
📋 Overview
Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know, no consolidated… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-cleaned.macedonian_sa
Sentiment Analysis Data for the Macedonian Language
Dataset Description:
This dataset contains a sentiment analysis dataset from Jovanski et al. (2015).
Data Structure:
The data was used for the project on improving word embeddings with graph knowledge for Low Resource Languages.
Citation:
@inproceedings{jovanoski-etal-2015-sentiment,
title = "Sentiment Analysis in {T}witter for {M}acedonian",
author = "Jovanoski, Dame and
Pachovski, Veno and
Nakov, Preslav"… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/macedonian_sa.macedonian-speech-datasetmacedonian-tweet-sentiment-classification
Dataset Card for Macedonian Tweet Sentiment Classification
Dataset Description
Dataset Summary
This is a Macedonian dataset is a collection of tweets for sentiment classification.
Supported Tasks and Leaderboards
text-classification, sentiment-classification: The dataset can be used to train a model for sentiment classification. The model performance is evaluated based on the accuracy of the predicted labels as compared to the given labels in the… See the full description on the dataset page: https://huggingface.co/datasets/isaacchung/macedonian-tweet-sentiment-classification.alpaca_macedonian_tacoThis repository contains the dataset used for the TaCo paper.
The dataset follows the style outlined in the TaCo paper, as follows:
{
"instruction": "instruction in xx",
"input": "input in xx",
"output": "Instruction in English: instruction in en ,
Response in English: response in en ,
Response in xx: response in xx "
}
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_macedonian_taco.alpaca-macedonian-cleanedThis repository contains the dataset used for the TaCo paper.
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation
@inproceedings{upadhayay2024taco,
title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes},
author={Bibek Upadhayay and Vahid Behzadan},
booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-macedonian-cleaned.Macedonian-Speech-Dataset
🎧 Macedonian Speech Dataset
The Macedonian Speech Dataset is a high-quality speech audio dataset designed to provide structured and reliable audio data for AI and machine learning systems. It includes 128 hours of audio data distributed across 654 files, delivered in MP3 and WAV formats, with a total size of 279 MB. This carefully curated audio dataset ensures diverse and representative voice data, with 55% female and 45% male speakers, and an age range spanning from 18 to 50+… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Macedonian-Speech-Dataset.macedonian-instructionsmacedonian_tweet_sentimentencyclopaedia-macedonicaexoDatamaceoff_r6_l1_128ch
