datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sklep
Dataset Card for skLEP
Dataset Description
skLEP (General Language Understanding Evaluation benchmark for Slovak) is the first comprehensive benchmark specifically designed for evaluating Slovak natural language understanding (NLU) models. The benchmark encompasses nine diverse tasks that span token-level, sentence-pair, and document-level challenges, thereby offering a thorough assessment of model capabilities.
To create this benchmark, we curated new, original datasets… See the full description on the dataset page: https://huggingface.co/datasets/slovak-nlp/sklep.fotkyzadarmo_slovak
FotkyZadarmo Slovak
A Slovak image-tagging dataset built from fotkyzadarmo.sk,
a CC0-licensed Slovak stock photo site.
Contents
Each row is one photo with its Slovak metadata:
image: the photo
tags: list of Slovak tags describing the image
title: the original Slovak caption
source_url: the original fotkyzadarmo.sk permalink, for provenance
~1,612 images.
Construction
Scraped from fotkyzadarmo.sk: title, tags, full-resolution image URL.
Tags are… See the full description on the dataset page: https://huggingface.co/datasets/driesaster/fotkyzadarmo_slovak.ipfs_slovakia_laws_ir
Slovakia legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_slovakia_laws (revision 22508a5eab4b3c0ba98fe0ab32aa9fcb5ec4be9a) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Slovakia prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_slovakia_laws_ir.fineweb2-slovak
FineWeb2 Slovak
This is the Slovak Portion of The FineWeb2 Dataset.
Known within subsets as slk_Latn, this language boasts an extensive corpus of over 14.1 billion words across more than 26.5 million documents.
Purpose of This Repository
This repository provides easy access to the Slovak portion of the extensive FineWeb2 dataset.
The existing dataset was extended with additional information, especially language identification using FastText, langdetect and lingua using… See the full description on the dataset page: https://huggingface.co/datasets/ivykopal/fineweb2-slovak.hate_speech_slovak
Slovak Hate Speech and Offensive Language Database
The dataset contains posts from a social network with human annotations.
Annotations
The posts are marked 1 if the post contain hateful or offensive language, 0 otherwise.
Dataset Creation
The source data were scraped from a social network from a selection of public pages for sport, politics or general discussion. The gathered data were cleaned from span with a text clustering.
The posts were annotated by a… See the full description on the dataset page: https://huggingface.co/datasets/TUKE-KEMT/hate_speech_slovak.SlovakMovieReviewSentimentClassification
SlovakMovieReviewSentimentClassification
An MTEB dataset
Massive Text Embedding Benchmark
User reviews of movies on the CSFD movie database, with 2 sentiment classes (positive, negative)
Task category
t2c
Domains
Reviews, Written
Reference
https://arxiv.org/pdf/2304.01922
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_task("SlovakMovieReviewSentimentClassification")… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SlovakMovieReviewSentimentClassification.slovakspeech_female_dataset
SlovakSpeechFemale TTS Dataset
suitable for TTS (NOT ASR)
slovak transcript is provided (NOT IPA)
48000 Hz sample rate
about 1 hour of audio
alpaca-slovak-cleanedThis repository contains the dataset used for the TaCo paper.
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation
@inproceedings{upadhayay2024taco,
title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes},
author={Bibek Upadhayay and Vahid Behzadan},
booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-slovak-cleaned.slovak-sft
Slovak SFT Dataset
A supervised fine-tuning (SFT) dataset for Slovak language instruction following, constructed from two publicly available Slovak resources:
saillab/alpaca-slovak-cleaned — Slovak instruction-response pairs
TUKE-DeutscheTelekom/skquad — Slovak question answering, rewritten into chat-style prompts
Format
Each example follows the standard messages format with three turns:
{
"messages": [
{"role": "system", "content": "Si užitočný slovenský… See the full description on the dataset page: https://huggingface.co/datasets/mbenco/slovak-sft.slovak-web-qa-pairsipfs_slovakia_laws
Slovakia Collection of Laws (SLOV-LEX)
Research snapshot of official national legislation from SLOV-LEX (Ministry of Justice) Collection of Laws static mirror.
Not legal advice. The official gazette / authentic source prevails over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-02
Coverage
complete
Source
SLOV-LEX (Ministry of Justice) Collection of Laws static mirror
Collector
scrapers/collect_sk.py
Laws / instruments
16,799
Articles… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_slovakia_laws.SlovakCOPA
Dataset Card for SlovakCOPA
Dataset Description
SlovakCOPA is a Slovak translation of the COPA (Choices of Plausible Alternatives) dataset, a benchmark for commonsense causal reasoning. COPA is a causal reasoning dataset where given a premise, the model must choose between two alternatives that are either the cause or the effect of the premise.
This dataset contains 600 examples (500 test / 100 validation) with parallel translations in English and standard Slovak.… See the full description on the dataset page: https://huggingface.co/datasets/NaiveNeuron/SlovakCOPA.slovak-triplets
Slovak Triplets Dataset
This repository contains the Slovak Triplets Dataset, a collection of triplet sentences in the Slovak language designed for training and evaluating document embedding models. Each triplet consists of an anchor sentence, a positive sentence (similar to the anchor), and a negative sentence (dissimilar to the anchor).
Data Source
The dataset is extracted form the Slovak part of WebFAQ and MQA datasets, which are publicly available collections of… See the full description on the dataset page: https://huggingface.co/datasets/TUKE-KEMT/slovak-triplets.UNER_Slovak-SNK
UNER_Slovak-SNK
The UNER dataset for the Slovak National Corpus (SNK), originally released with
UNER v1.
UNER_Slovak-SNK is part of Universal NER and is based on the UD_Slovak-SNK dataset.
The canonical reference commit to the Universal Dependencies dataset is 4d13f4810ebba49ed41430c3e43adf5a50087f6f
If you use this dataset, please cite the corresponding paper:
@inproceedings{
mayhew2024universal,
title={Universal NER: A Gold-Standard Multilingual Named Entity Recognition… See the full description on the dataset page: https://huggingface.co/datasets/universalner/UNER_Slovak-SNK.gigatrue-slovak
Gigatrue Slovak abstractive summarisation dataset.
Synthetic Gigaword dataset translated to Slovak.
Same as https://huggingface.co/datasets/Plasmoxy/gigatrue but translated to Slovak using SeamlessM4T-v2 (https://huggingface.co/docs/transformers/en/model_doc/seamless_m4t_v2).
Original dataset adapted from https://huggingface.co/datasets/Harvard/gigaword.
This work is supported by the EU NextGenerationEU through the Recovery and Resilience Plan for Slovakia under the project No.… See the full description on the dataset page: https://huggingface.co/datasets/Plasmoxy/gigatrue-slovak.SlovakSTS
SlovakSTS
An MTEB dataset
Massive Text Embedding Benchmark
A professional Slovak translation of the STS Benchmark (STSb), originally part of the GLUE benchmark. The task is Semantic Textual Similarity (STS): given a pair of sentences, the goal is to predict their semantic similarity on a continuous scale from 0 (completely unrelated) to 5 (semantically equivalent). Sentence pairs are drawn from news headlines, image captions, and forum posts.
Task category
STS… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SlovakSTS.slovaksumThe SlovakSum dataset from the SlovakSum: Slovak News Summarization Dataset paper
jz_sync_Open-Government-Slovak-filteredtree_species_classification_slovak_normal_cropped
Tree Species Classification Slovak Normal Cropped
This dataset provides real RGB images of tree species collected in natural field environments across Slovakia and the Czech Republic. Images were captured using handheld cameras (Sony Alpha 7 and Canon EOS 4000D) during field surveys in June and August 2022. The dataset contains 1,367 images across 4 classes: European beech, European silver fir, Norway spruce, Sessile oak.Images per class:
European beech: 300
European silver… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/tree_species_classification_slovak_normal_cropped.slovak-sft
Slovak SFT Dataset
A supervised fine-tuning (SFT) dataset for Slovak language instruction following, constructed from two publicly available Slovak resources:
saillab/alpaca-slovak-cleaned — Slovak instruction-response pairs
TUKE-DeutscheTelekom/skquad — Slovak question answering, rewritten into chat-style prompts
Format
Each example follows the standard messages format with three turns:
{
"messages": [
{"role": "system", "content": "Si užitočný slovenský… See the full description on the dataset page: https://huggingface.co/datasets/Adamos3000/slovak-sft.tree_species_classification_slovak_exact_cropped
Tree Species Classification Slovak Exact Cropped
This dataset provides real RGB images of tree bark and foliage collected in field environments across Slovakia and the Czech Republic. Captured during summer 2022 using handheld cameras (Sony Alpha 7 and Canon EOS 4000D), the images focus on multiple tree species for forestry classification tasks. The dataset contains 527 images across 4 classes: European beech, European silver fir, Norway spruce, Sessile oak.Images per class:… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/tree_species_classification_slovak_exact_cropped.alpaca_slovak_tacoThis repository contains the dataset used for the TaCo paper.
The dataset follows the style outlined in the TaCo paper, as follows:
{
"instruction": "instruction in xx",
"input": "input in xx",
"output": "Instruction in English: instruction in en ,
Response in English: response in en ,
Response in xx: response in xx "
}
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_slovak_taco.SlovakSumRetrieval
SlovakSumRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
SlovakSum, a Slovak news summarization dataset consisting of over 200 thousand
news articles with titles and short abstracts obtained from multiple Slovak newspapers.
Originally intended as a summarization task, but since no human annotations were provided
here reformulated to a retrieval task.
Task category
t2t
Domains
News, Social, Web, Written
Reference… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SlovakSumRetrieval.SlovakSumURLClustering
SlovakSumURLClustering
An MTEB dataset
Massive Text Embedding Benchmark
Clustering of Slovak news articles from SlovakSum dataset based on the URL structure. Articles are organized into 12 editorial categories including sports, culture, economy, health, travel, politics, and technology sections.
Task category
Clustering (text-to-category)
Domains
News, Written
Reference
Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SlovakSumURLClustering.tree_species_classification_slovak_exact_cropped_extended
Tree Species Classification Slovak Exact Cropped Extended
This dataset provides real RGB images of tree species collected in natural field environments across Slovakia and the Czech Republic. Images were captured using handheld cameras (Sony Alpha 7 and Canon EOS 4000D) during field campaigns in June and August 2022, focusing on bark and foliage for species identification. The dataset contains 1,367 images across 4 classes: European beech, European silver fir, Norway spruce… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/tree_species_classification_slovak_exact_cropped_extended.exam-slovak-mathbioPart of INCLUDE
see research: https://arxiv.org/abs/2411.19799
full dataset: https://huggingface.co/datasets/CohereForAI/include-base-44
slovakspeech_male_dataset
SlovakSpeechMale TTS Dataset
suitable for TTS (NOT ASR)
slovak transcript is provided (NOT IPA)
48000 Hz sample rate
about 1 hour of audio
slovak-speech-datasetslovak-pharmacy-drmax-reranking
SlovakPharmacyDrMaxReranking
An MTEB dataset
Massive Text Embedding Benchmark
A reranking dataset created from Q&A content collected from DrMax pharmacy website. The dataset consists of questions about medications, health conditions, and pharmaceutical advice, with answers provided by qualified pharmacists. This dataset is designed to evaluate models' ability to rank relevant pharmaceutical information and expert responses.
Task category
t2t
Domains
Medical, Web… See the full description on the dataset page: https://huggingface.co/datasets/slovak-nlp/slovak-pharmacy-drmax-reranking.slovak-scrape
