datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hate_speech_slovak
Slovak Hate Speech and Offensive Language Database
The dataset contains posts from a social network with human annotations.
Annotations
The posts are marked 1 if the post contain hateful or offensive language, 0 otherwise.
Dataset Creation
The source data were scraped from a social network from a selection of public pages for sport, politics or general discussion. The gathered data were cleaned from span with a text clustering.
The posts were annotated by a… See the full description on the dataset page: https://huggingface.co/datasets/TUKE-KEMT/hate_speech_slovak.SlovakBebrasChallenge
iBobor Slovak Questions Dataset
Dataset Description
This dataset contains questions from the Slovak iBobor competition, publicly available on the website http://demo.ibobor.sk/sutaz_demo/, which is part of the international Bebras Challenge focused on informatics and computational thinking.
The dataset was created for the purposes of a master's thesis at Comenius University in Bratislava, Faculty of Mathematics, Physics and Informatics. The thesis is concerned with… See the full description on the dataset page: https://huggingface.co/datasets/patriciavnencakova/SlovakBebrasChallenge.slovak-financial-exam
Dataset Card for Slovak Financial QA
Dataset Description
This dataset contains 1,334 multiple-choice questions from the financial domain in the Slovak language. It was created to address the limited availability of language resources for Slovak, providing a benchmark for evaluating language models' capabilities in a specialized, low-resource domain.
The questions are sourced from the official certification exams for financial advisors in Slovakia, covering a range of… See the full description on the dataset page: https://huggingface.co/datasets/TUKE-KEMT/slovak-financial-exam.
