Slovak
Datasets
All datasets matching “Slovak”sklep
Dataset Card for skLEP
Dataset Description
skLEP (General Language Understanding Evaluation benchmark for Slovak) is the first comprehensive benchmark specifically designed for evaluating Slovak natural language understanding (NLU) models. The benchmark encompasses nine diverse tasks that span token-level, sentence-pair, and document-level challenges, thereby offering a thorough assessment of model capabilities.
To create this benchmark, we curated new, original datasets… See the full description on the dataset page: https://huggingface.co/datasets/slovak-nlp/sklep.fotkyzadarmo_slovak
FotkyZadarmo Slovak
A Slovak image-tagging dataset built from fotkyzadarmo.sk,
a CC0-licensed Slovak stock photo site.
Contents
Each row is one photo with its Slovak metadata:
image: the photo
tags: list of Slovak tags describing the image
title: the original Slovak caption
source_url: the original fotkyzadarmo.sk permalink, for provenance
~1,612 images.
Construction
Scraped from fotkyzadarmo.sk: title, tags, full-resolution image URL.
Tags are… See the full description on the dataset page: https://huggingface.co/datasets/driesaster/fotkyzadarmo_slovak.ipfs_slovakia_laws_ir
Slovakia legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_slovakia_laws (revision 22508a5eab4b3c0ba98fe0ab32aa9fcb5ec4be9a) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Slovakia prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_slovakia_laws_ir.hate_speech_slovak
Slovak Hate Speech and Offensive Language Database
The dataset contains posts from a social network with human annotations.
Annotations
The posts are marked 1 if the post contain hateful or offensive language, 0 otherwise.
Dataset Creation
The source data were scraped from a social network from a selection of public pages for sport, politics or general discussion. The gathered data were cleaned from span with a text clustering.
The posts were annotated by a… See the full description on the dataset page: https://huggingface.co/datasets/TUKE-KEMT/hate_speech_slovak.fineweb2-slovak
FineWeb2 Slovak
This is the Slovak Portion of The FineWeb2 Dataset.
Known within subsets as slk_Latn, this language boasts an extensive corpus of over 14.1 billion words across more than 26.5 million documents.
Purpose of This Repository
This repository provides easy access to the Slovak portion of the extensive FineWeb2 dataset.
The existing dataset was extended with additional information, especially language identification using FastText, langdetect and lingua using… See the full description on the dataset page: https://huggingface.co/datasets/ivykopal/fineweb2-slovak.SlovakMovieReviewSentimentClassification
SlovakMovieReviewSentimentClassification
An MTEB dataset
Massive Text Embedding Benchmark
User reviews of movies on the CSFD movie database, with 2 sentiment classes (positive, negative)
Task category
t2c
Domains
Reviews, Written
Reference
https://arxiv.org/pdf/2304.01922
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_task("SlovakMovieReviewSentimentClassification")… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SlovakMovieReviewSentimentClassification.
