CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01slovak-nlp /sklep Dataset Card for skLEP Dataset Description skLEP (General Language Understanding Evaluation benchmark for Slovak) is the first comprehensive benchmark specifically designed for evaluating Slovak natural language understanding (NLU) models. The benchmark encompasses nine diverse tasks that span token-level, sentence-pair, and document-level challenges, thereby offering a thorough assessment of model capabilities. To create this benchmark, we curated new, original datasets… See the full description on the dataset page: https://huggingface.co/datasets/slovak-nlp/sklep.textquestion-answering100K<n<1M4 likes437 downloads7mo agoHugging Face02driesaster /fotkyzadarmo_slovak FotkyZadarmo Slovak A Slovak image-tagging dataset built from fotkyzadarmo.sk, a CC0-licensed Slovak stock photo site. Contents Each row is one photo with its Slovak metadata: image: the photo tags: list of Slovak tags describing the image title: the original Slovak caption source_url: the original fotkyzadarmo.sk permalink, for provenance ~1,612 images. Construction Scraped from fotkyzadarmo.sk: title, tags, full-resolution image URL. Tags are… See the full description on the dataset page: https://huggingface.co/datasets/driesaster/fotkyzadarmo_slovak.imagevisual-question-answering1K<n<10K0 likes248 downloads1mo agoHugging Face03justicedao /ipfs_slovakia_laws_ir Slovakia legislation IR (CID-keyed sparse GraphRAG) Research retrieval release of endomorphosis/ipfs_slovakia_laws (revision 22508a5eab4b3c0ba98fe0ab32aa9fcb5ec4be9a) packaged as country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir). Not legal advice. This is a research snapshot. The official gazette / authentic source of Slovakia prevails over this corpus. Retrieved documents and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_slovakia_laws_ir.tabulartext-retrieval100K<n<1M0 likes228 downloads2d agoHugging Face04ivykopal /fineweb2-slovak FineWeb2 Slovak This is the Slovak Portion of The FineWeb2 Dataset. Known within subsets as slk_Latn, this language boasts an extensive corpus of over 14.1 billion words across more than 26.5 million documents. Purpose of This Repository This repository provides easy access to the Slovak portion of the extensive FineWeb2 dataset. The existing dataset was extended with additional information, especially language identification using FastText, langdetect and lingua using… See the full description on the dataset page: https://huggingface.co/datasets/ivykopal/fineweb2-slovak.tabular10M<n<100M3 likes212 downloads1y agoHugging Face05TUKE-KEMT /hate_speech_slovak Slovak Hate Speech and Offensive Language Database The dataset contains posts from a social network with human annotations. Annotations The posts are marked 1 if the post contain hateful or offensive language, 0 otherwise. Dataset Creation The source data were scraped from a social network from a selection of public pages for sport, politics or general discussion. The gathered data were cleaned from span with a text clustering. The posts were annotated by a… See the full description on the dataset page: https://huggingface.co/datasets/TUKE-KEMT/hate_speech_slovak.tabulartext-classification10K<n<100K5 likes210 downloads2y agoHugging Face06mteb /SlovakMovieReviewSentimentClassification SlovakMovieReviewSentimentClassification An MTEB dataset Massive Text Embedding Benchmark User reviews of movies on the CSFD movie database, with 2 sentiment classes (positive, negative) Task category t2c Domains Reviews, Written Reference https://arxiv.org/pdf/2304.01922 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("SlovakMovieReviewSentimentClassification")… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SlovakMovieReviewSentimentClassification.texttext-classification10K<n<100K0 likes125 downloads1y agoHugging Face07neurlang /slovakspeech_female_dataset SlovakSpeechFemale TTS Dataset suitable for TTS (NOT ASR) slovak transcript is provided (NOT IPA) 48000 Hz sample rate about 1 hour of audio 2 likes107 downloads6mo agoHugging Face08saillab /alpaca-slovak-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-slovak-cleaned.text10K<n<100K0 likes81 downloads2y agoHugging Face09mbenco /slovak-sft Slovak SFT Dataset A supervised fine-tuning (SFT) dataset for Slovak language instruction following, constructed from two publicly available Slovak resources: saillab/alpaca-slovak-cleaned — Slovak instruction-response pairs TUKE-DeutscheTelekom/skquad — Slovak question answering, rewritten into chat-style prompts Format Each example follows the standard messages format with three turns: { "messages": [ {"role": "system", "content": "Si užitočný slovenský… See the full description on the dataset page: https://huggingface.co/datasets/mbenco/slovak-sft.texttext-generation10K<n<100K0 likes81 downloads5mo agoHugging Face10TUKE-KEMT /slovak-web-qa-pairstext100K<n<1M0 likes79 downloads9mo agoHugging Face11endomorphosis /ipfs_slovakia_laws Slovakia Collection of Laws (SLOV-LEX) Research snapshot of official national legislation from SLOV-LEX (Ministry of Justice) Collection of Laws static mirror. Not legal advice. The official gazette / authentic source prevails over this corpus. Snapshot Field Value Snapshot date 2026-09-02 Coverage complete Source SLOV-LEX (Ministry of Justice) Collection of Laws static mirror Collector scrapers/collect_sk.py Laws / instruments 16,799 Articles… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_slovakia_laws.texttext-retrieval10K<n<100K0 likes78 downloads22d agoHugging Face12NaiveNeuron /SlovakCOPA Dataset Card for SlovakCOPA Dataset Description SlovakCOPA is a Slovak translation of the COPA (Choices of Plausible Alternatives) dataset, a benchmark for commonsense causal reasoning. COPA is a causal reasoning dataset where given a premise, the model must choose between two alternatives that are either the cause or the effect of the premise. This dataset contains 600 examples (500 test / 100 validation) with parallel translations in English and standard Slovak.… See the full description on the dataset page: https://huggingface.co/datasets/NaiveNeuron/SlovakCOPA.tabularmultiple-choicen<1K1 likes69 downloads5mo agoHugging Face13TUKE-KEMT /slovak-triplets Slovak Triplets Dataset This repository contains the Slovak Triplets Dataset, a collection of triplet sentences in the Slovak language designed for training and evaluating document embedding models. Each triplet consists of an anchor sentence, a positive sentence (similar to the anchor), and a negative sentence (dissimilar to the anchor). Data Source The dataset is extracted form the Slovak part of WebFAQ and MQA datasets, which are publicly available collections of… See the full description on the dataset page: https://huggingface.co/datasets/TUKE-KEMT/slovak-triplets.text100K<n<1M1 likes59 downloads8mo agoHugging Face14universalner /UNER_Slovak-SNK UNER_Slovak-SNK The UNER dataset for the Slovak National Corpus (SNK), originally released with UNER v1. UNER_Slovak-SNK is part of Universal NER and is based on the UD_Slovak-SNK dataset. The canonical reference commit to the Universal Dependencies dataset is 4d13f4810ebba49ed41430c3e43adf5a50087f6f If you use this dataset, please cite the corresponding paper: @inproceedings{ mayhew2024universal, title={Universal NER: A Gold-Standard Multilingual Named Entity Recognition… See the full description on the dataset page: https://huggingface.co/datasets/universalner/UNER_Slovak-SNK.0 likes58 downloads2y agoHugging Face15Plasmoxy /gigatrue-slovak Gigatrue Slovak abstractive summarisation dataset. Synthetic Gigaword dataset translated to Slovak. Same as https://huggingface.co/datasets/Plasmoxy/gigatrue but translated to Slovak using SeamlessM4T-v2 (https://huggingface.co/docs/transformers/en/model_doc/seamless_m4t_v2). Original dataset adapted from https://huggingface.co/datasets/Harvard/gigaword. This work is supported by the EU NextGenerationEU through the Recovery and Resilience Plan for Slovakia under the project No.… See the full description on the dataset page: https://huggingface.co/datasets/Plasmoxy/gigatrue-slovak.textsummarization1M<n<10M2 likes58 downloads2y agoHugging Face16mteb /SlovakSTS SlovakSTS An MTEB dataset Massive Text Embedding Benchmark A professional Slovak translation of the STS Benchmark (STSb), originally part of the GLUE benchmark. The task is Semantic Textual Similarity (STS): given a pair of sentences, the goal is to predict their semantic similarity on a continuous scale from 0 (completely unrelated) to 5 (semantically equivalent). Sentence pairs are drawn from news headlines, image captions, and forum posts. Task category STS… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SlovakSTS.textsentence-similarity1K<n<10K0 likes54 downloads1mo agoHugging Face17NaiveNeuron /slovaksumThe SlovakSum dataset from the SlovakSum: Slovak News Summarization Dataset paper text100K<n<1M8 likes50 downloads3y agoHugging Face18JZSG /jz_sync_Open-Government-Slovak-filteredtabular100K<n<1M0 likes39 downloads11mo agoHugging Face19Project-AgML /tree_species_classification_slovak_normal_cropped Tree Species Classification Slovak Normal Cropped This dataset provides real RGB images of tree species collected in natural field environments across Slovakia and the Czech Republic. Images were captured using handheld cameras (Sony Alpha 7 and Canon EOS 4000D) during field surveys in June and August 2022. The dataset contains 1,367 images across 4 classes: European beech, European silver fir, Norway spruce, Sessile oak.Images per class: European beech: 300 European silver… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/tree_species_classification_slovak_normal_cropped.imageimage-classification1K<n<10K0 likes33 downloads19h agoHugging Face20Adamos3000 /slovak-sft Slovak SFT Dataset A supervised fine-tuning (SFT) dataset for Slovak language instruction following, constructed from two publicly available Slovak resources: saillab/alpaca-slovak-cleaned — Slovak instruction-response pairs TUKE-DeutscheTelekom/skquad — Slovak question answering, rewritten into chat-style prompts Format Each example follows the standard messages format with three turns: { "messages": [ {"role": "system", "content": "Si užitočný slovenský… See the full description on the dataset page: https://huggingface.co/datasets/Adamos3000/slovak-sft.texttext-generation10K<n<100K0 likes32 downloads3mo agoHugging Face21Project-AgML /tree_species_classification_slovak_exact_cropped Tree Species Classification Slovak Exact Cropped This dataset provides real RGB images of tree bark and foliage collected in field environments across Slovakia and the Czech Republic. Captured during summer 2022 using handheld cameras (Sony Alpha 7 and Canon EOS 4000D), the images focus on multiple tree species for forestry classification tasks. The dataset contains 527 images across 4 classes: European beech, European silver fir, Norway spruce, Sessile oak.Images per class:… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/tree_species_classification_slovak_exact_cropped.imageimage-classificationn<1K0 likes32 downloads19h agoHugging Face22saillab /alpaca_slovak_tacoThis repository contains the dataset used for the TaCo paper. The dataset follows the style outlined in the TaCo paper, as follows: { "instruction": "instruction in xx", "input": "input in xx", "output": "Instruction in English: instruction in en , Response in English: response in en , Response in xx: response in xx " } Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_slovak_taco.text10K<n<100K0 likes31 downloads2y agoHugging Face23mteb /SlovakSumRetrieval SlovakSumRetrieval An MTEB dataset Massive Text Embedding Benchmark SlovakSum, a Slovak news summarization dataset consisting of over 200 thousand news articles with titles and short abstracts obtained from multiple Slovak newspapers. Originally intended as a summarization task, but since no human annotations were provided here reformulated to a retrieval task. Task category t2t Domains News, Social, Web, Written Reference… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SlovakSumRetrieval.texttext-retrieval1K<n<10K0 likes31 downloads1y agoHugging Face24mteb /SlovakSumURLClustering SlovakSumURLClustering An MTEB dataset Massive Text Embedding Benchmark Clustering of Slovak news articles from SlovakSum dataset based on the URL structure. Articles are organized into 12 editorial categories including sports, culture, economy, health, travel, politics, and technology sections. Task category Clustering (text-to-category) Domains News, Written Reference Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SlovakSumURLClustering.texttext-classification10K<n<100K0 likes29 downloads1mo agoHugging Face25Project-AgML /tree_species_classification_slovak_exact_cropped_extended Tree Species Classification Slovak Exact Cropped Extended This dataset provides real RGB images of tree species collected in natural field environments across Slovakia and the Czech Republic. Images were captured using handheld cameras (Sony Alpha 7 and Canon EOS 4000D) during field campaigns in June and August 2022, focusing on bark and foliage for species identification. The dataset contains 1,367 images across 4 classes: European beech, European silver fir, Norway spruce… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/tree_species_classification_slovak_exact_cropped_extended.imageimage-classification1K<n<10K0 likes29 downloads19h agoHugging Face26dokato /exam-slovak-mathbioPart of INCLUDE see research: https://arxiv.org/abs/2411.19799 full dataset: https://huggingface.co/datasets/CohereForAI/include-base-44 textmultiple-choicen<1K1 likes28 downloads2y agoHugging Face27neurlang /slovakspeech_male_dataset SlovakSpeechMale TTS Dataset suitable for TTS (NOT ASR) slovak transcript is provided (NOT IPA) 48000 Hz sample rate about 1 hour of audio 2 likes28 downloads4mo agoHugging Face28shunyalabs /slovak-speech-datasetaudio1K<n<10K0 likes27 downloads1y agoHugging Face29slovak-nlp /slovak-pharmacy-drmax-rerankinggated SlovakPharmacyDrMaxReranking An MTEB dataset Massive Text Embedding Benchmark A reranking dataset created from Q&A content collected from DrMax pharmacy website. The dataset consists of questions about medications, health conditions, and pharmaceutical advice, with answers provided by qualified pharmacists. This dataset is designed to evaluate models' ability to rank relevant pharmaceutical information and expert responses. Task category t2t Domains Medical, Web… See the full description on the dataset page: https://huggingface.co/datasets/slovak-nlp/slovak-pharmacy-drmax-reranking.texttext-ranking10K<n<100K0 likes27 downloads4mo agoHugging Face30simonko912 /slovak-scrapetext1K<n<10K1 likes26 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.