datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fotkyzadarmo_slovak
FotkyZadarmo Slovak
A Slovak image-tagging dataset built from fotkyzadarmo.sk,
a CC0-licensed Slovak stock photo site.
Contents
Each row is one photo with its Slovak metadata:
image: the photo
tags: list of Slovak tags describing the image
title: the original Slovak caption
source_url: the original fotkyzadarmo.sk permalink, for provenance
~1,612 images.
Construction
Scraped from fotkyzadarmo.sk: title, tags, full-resolution image URL.
Tags are… See the full description on the dataset page: https://huggingface.co/datasets/driesaster/fotkyzadarmo_slovak.sklep
Dataset Card for skLEP
Dataset Description
skLEP (General Language Understanding Evaluation benchmark for Slovak) is the first comprehensive benchmark specifically designed for evaluating Slovak natural language understanding (NLU) models. The benchmark encompasses nine diverse tasks that span token-level, sentence-pair, and document-level challenges, thereby offering a thorough assessment of model capabilities.
To create this benchmark, we curated new, original datasets… See the full description on the dataset page: https://huggingface.co/datasets/slovak-nlp/sklep.ipfs_slovakia_laws_ir
justicedao/ipfs_slovakia_laws_ir
Slovakia laws IR release.
fineweb2-slovak
FineWeb2 Slovak
This is the Slovak Portion of The FineWeb2 Dataset.
Known within subsets as slk_Latn, this language boasts an extensive corpus of over 14.1 billion words across more than 26.5 million documents.
Purpose of This Repository
This repository provides easy access to the Slovak portion of the extensive FineWeb2 dataset.
The existing dataset was extended with additional information, especially language identification using FastText, langdetect and lingua using… See the full description on the dataset page: https://huggingface.co/datasets/ivykopal/fineweb2-slovak.hate_speech_slovak
Slovak Hate Speech and Offensive Language Database
The dataset contains posts from a social network with human annotations.
Annotations
The posts are marked 1 if the post contain hateful or offensive language, 0 otherwise.
Dataset Creation
The source data were scraped from a social network from a selection of public pages for sport, politics or general discussion. The gathered data were cleaned from span with a text clustering.
The posts were annotated by a… See the full description on the dataset page: https://huggingface.co/datasets/TUKE-KEMT/hate_speech_slovak.SlovakMovieReviewSentimentClassification
SlovakMovieReviewSentimentClassification
An MTEB dataset
Massive Text Embedding Benchmark
User reviews of movies on the CSFD movie database, with 2 sentiment classes (positive, negative)
Task category
t2c
Domains
Reviews, Written
Reference
https://arxiv.org/pdf/2304.01922
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_task("SlovakMovieReviewSentimentClassification")… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SlovakMovieReviewSentimentClassification.ipfs_slovakia_laws
Slovakia Collection of Laws (SLOV-LEX)
Research snapshot of official national legislation from SLOV-LEX (Ministry of Justice) Collection of Laws static mirror.
Not legal advice. The official gazette / authentic source prevails over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-02
Coverage
complete
Source
SLOV-LEX (Ministry of Justice) Collection of Laws static mirror
Collector
scrapers/collect_sk.py
Laws / instruments
16,799
Articles… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_slovakia_laws.alpaca-slovak-cleanedThis repository contains the dataset used for the TaCo paper.
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation
@inproceedings{upadhayay2024taco,
title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes},
author={Bibek Upadhayay and Vahid Behzadan},
booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-slovak-cleaned.slovak-sft
Slovak SFT Dataset
A supervised fine-tuning (SFT) dataset for Slovak language instruction following, constructed from two publicly available Slovak resources:
saillab/alpaca-slovak-cleaned — Slovak instruction-response pairs
TUKE-DeutscheTelekom/skquad — Slovak question answering, rewritten into chat-style prompts
Format
Each example follows the standard messages format with three turns:
{
"messages": [
{"role": "system", "content": "Si užitočný slovenský… See the full description on the dataset page: https://huggingface.co/datasets/mbenco/slovak-sft.slovak-web-qa-pairsSlovakCOPA
Dataset Card for SlovakCOPA
Dataset Description
SlovakCOPA is a Slovak translation of the COPA (Choices of Plausible Alternatives) dataset, a benchmark for commonsense causal reasoning. COPA is a causal reasoning dataset where given a premise, the model must choose between two alternatives that are either the cause or the effect of the premise.
This dataset contains 600 examples (500 test / 100 validation) with parallel translations in English and standard Slovak.… See the full description on the dataset page: https://huggingface.co/datasets/NaiveNeuron/SlovakCOPA.gigatrue-slovak
Gigatrue Slovak abstractive summarisation dataset.
Synthetic Gigaword dataset translated to Slovak.
Same as https://huggingface.co/datasets/Plasmoxy/gigatrue but translated to Slovak using SeamlessM4T-v2 (https://huggingface.co/docs/transformers/en/model_doc/seamless_m4t_v2).
Original dataset adapted from https://huggingface.co/datasets/Harvard/gigaword.
This work is supported by the EU NextGenerationEU through the Recovery and Resilience Plan for Slovakia under the project No.… See the full description on the dataset page: https://huggingface.co/datasets/Plasmoxy/gigatrue-slovak.SlovakSTS
SlovakSTS
An MTEB dataset
Massive Text Embedding Benchmark
A professional Slovak translation of the STS Benchmark (STSb), originally part of the GLUE benchmark. The task is Semantic Textual Similarity (STS): given a pair of sentences, the goal is to predict their semantic similarity on a continuous scale from 0 (completely unrelated) to 5 (semantically equivalent). Sentence pairs are drawn from news headlines, image captions, and forum posts.
Task category
STS… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SlovakSTS.slovaksumThe SlovakSum dataset from the SlovakSum: Slovak News Summarization Dataset paper
slovak-triplets
Slovak Triplets Dataset
This repository contains the Slovak Triplets Dataset, a collection of triplet sentences in the Slovak language designed for training and evaluating document embedding models. Each triplet consists of an anchor sentence, a positive sentence (similar to the anchor), and a negative sentence (dissimilar to the anchor).
Data Source
The dataset is extracted form the Slovak part of WebFAQ and MQA datasets, which are publicly available collections of… See the full description on the dataset page: https://huggingface.co/datasets/TUKE-KEMT/slovak-triplets.jz_sync_Open-Government-Slovak-filteredtree_species_classification_slovak_exact_cropped_extended
Tree Species Classification Slovak Exact Cropped Extended
This dataset provides real RGB images of tree species collected in natural field environments across Slovakia and the Czech Republic. Images were captured using handheld cameras (Sony Alpha 7 and Canon EOS 4000D) during field campaigns in June and August 2022, focusing on bark and foliage for species identification. The dataset contains 1,367 images across 4 classes: European beech, European silver fir, Norway spruce… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/tree_species_classification_slovak_exact_cropped_extended.slovak-sft
Slovak SFT Dataset
A supervised fine-tuning (SFT) dataset for Slovak language instruction following, constructed from two publicly available Slovak resources:
saillab/alpaca-slovak-cleaned — Slovak instruction-response pairs
TUKE-DeutscheTelekom/skquad — Slovak question answering, rewritten into chat-style prompts
Format
Each example follows the standard messages format with three turns:
{
"messages": [
{"role": "system", "content": "Si užitočný slovenský… See the full description on the dataset page: https://huggingface.co/datasets/Adamos3000/slovak-sft.SlovakSumRetrieval
SlovakSumRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
SlovakSum, a Slovak news summarization dataset consisting of over 200 thousand
news articles with titles and short abstracts obtained from multiple Slovak newspapers.
Originally intended as a summarization task, but since no human annotations were provided
here reformulated to a retrieval task.
Task category
t2t
Domains
News, Social, Web, Written
Reference… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SlovakSumRetrieval.SlovakSumURLClustering
SlovakSumURLClustering
An MTEB dataset
Massive Text Embedding Benchmark
Clustering of Slovak news articles from SlovakSum dataset based on the URL structure. Articles are organized into 12 editorial categories including sports, culture, economy, health, travel, politics, and technology sections.
Task category
Clustering (text-to-category)
Domains
News, Written
Reference
Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SlovakSumURLClustering.alpaca_slovak_tacoThis repository contains the dataset used for the TaCo paper.
The dataset follows the style outlined in the TaCo paper, as follows:
{
"instruction": "instruction in xx",
"input": "input in xx",
"output": "Instruction in English: instruction in en ,
Response in English: response in en ,
Response in xx: response in xx "
}
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_slovak_taco.slovak-speech-datasetexam-slovak-mathbioPart of INCLUDE
see research: https://arxiv.org/abs/2411.19799
full dataset: https://huggingface.co/datasets/CohereForAI/include-base-44
slovak-scrapeslovak-pharmacy-drmax-reranking
SlovakPharmacyDrMaxReranking
An MTEB dataset
Massive Text Embedding Benchmark
A reranking dataset created from Q&A content collected from DrMax pharmacy website. The dataset consists of questions about medications, health conditions, and pharmaceutical advice, with answers provided by qualified pharmacists. This dataset is designed to evaluate models' ability to rank relevant pharmaceutical information and expert responses.
Task category
t2t
Domains
Medical, Web… See the full description on the dataset page: https://huggingface.co/datasets/slovak-nlp/slovak-pharmacy-drmax-reranking.Slovak_Relation_Extractionslovak-pharmacy-mojalekaren-reranking
SlovakPharmacyMojaLekarenReranking
An MTEB dataset
Massive Text Embedding Benchmark
A reranking dataset created from Q&A content collected from MojaLekaren pharmacy website. The dataset consists of questions about medications, health conditions, and pharmaceutical advice, with answers provided by qualified pharmacists. This dataset is designed to evaluate models' ability to rank relevant pharmaceutical information and expert responses.
Task category
t2t
Domains
Medical… See the full description on the dataset page: https://huggingface.co/datasets/slovak-nlp/slovak-pharmacy-mojalekaren-reranking.slovak_sa
Sentiment Analysis Data for the Slovak Language
Dataset Description:
This dataset contains a sentiment analysis dataset from Pecar et al. (2019).
Data Structure:
The data was used for the project on improving word embeddings with graph knowledge for Low Resource Languages.
Citation:
@inproceedings{pecar-etal-2019-improving,
title = "Improving Sentiment Classification in {S}lovak Language",
author = "Pecar, Samuel and
Simko, Marian and
Bielikova, Maria"… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/slovak_sa.slovaksum-url-clusteringSlovakHateSpeechClassification
SlovakHateSpeechClassification
An MTEB dataset
Massive Text Embedding Benchmark
The dataset contains posts from a social network with human annotations for hateful or offensive language in Slovak.
Task category
t2c
Domains
Social, Written
Reference
https://huggingface.co/datasets/TUKE-KEMT/hate_speech_slovak
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SlovakHateSpeechClassification.
