CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01driesaster /fotkyzadarmo_slovak FotkyZadarmo Slovak A Slovak image-tagging dataset built from fotkyzadarmo.sk, a CC0-licensed Slovak stock photo site. Contents Each row is one photo with its Slovak metadata: image: the photo tags: list of Slovak tags describing the image title: the original Slovak caption source_url: the original fotkyzadarmo.sk permalink, for provenance ~1,612 images. Construction Scraped from fotkyzadarmo.sk: title, tags, full-resolution image URL. Tags are… See the full description on the dataset page: https://huggingface.co/datasets/driesaster/fotkyzadarmo_slovak.imagevisual-question-answering1K<n<10K0 likes462 downloads2mo agoHugging Face02slovak-nlp /sklep Dataset Card for skLEP Dataset Description skLEP (General Language Understanding Evaluation benchmark for Slovak) is the first comprehensive benchmark specifically designed for evaluating Slovak natural language understanding (NLU) models. The benchmark encompasses nine diverse tasks that span token-level, sentence-pair, and document-level challenges, thereby offering a thorough assessment of model capabilities. To create this benchmark, we curated new, original datasets… See the full description on the dataset page: https://huggingface.co/datasets/slovak-nlp/sklep.textquestion-answering100K<n<1M4 likes429 downloads7mo agoHugging Face03justicedao /ipfs_slovakia_laws_ir justicedao/ipfs_slovakia_laws_ir Slovakia laws IR release. tabular1M<n<10M0 likes245 downloads17h agoHugging Face04ivykopal /fineweb2-slovak FineWeb2 Slovak This is the Slovak Portion of The FineWeb2 Dataset. Known within subsets as slk_Latn, this language boasts an extensive corpus of over 14.1 billion words across more than 26.5 million documents. Purpose of This Repository This repository provides easy access to the Slovak portion of the extensive FineWeb2 dataset. The existing dataset was extended with additional information, especially language identification using FastText, langdetect and lingua using… See the full description on the dataset page: https://huggingface.co/datasets/ivykopal/fineweb2-slovak.tabular10M<n<100M3 likes213 downloads1y agoHugging Face05TUKE-KEMT /hate_speech_slovak Slovak Hate Speech and Offensive Language Database The dataset contains posts from a social network with human annotations. Annotations The posts are marked 1 if the post contain hateful or offensive language, 0 otherwise. Dataset Creation The source data were scraped from a social network from a selection of public pages for sport, politics or general discussion. The gathered data were cleaned from span with a text clustering. The posts were annotated by a… See the full description on the dataset page: https://huggingface.co/datasets/TUKE-KEMT/hate_speech_slovak.tabulartext-classification10K<n<100K5 likes208 downloads2y agoHugging Face06mteb /SlovakMovieReviewSentimentClassification SlovakMovieReviewSentimentClassification An MTEB dataset Massive Text Embedding Benchmark User reviews of movies on the CSFD movie database, with 2 sentiment classes (positive, negative) Task category t2c Domains Reviews, Written Reference https://arxiv.org/pdf/2304.01922 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("SlovakMovieReviewSentimentClassification")… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SlovakMovieReviewSentimentClassification.texttext-classification10K<n<100K0 likes126 downloads1y agoHugging Face07endomorphosis /ipfs_slovakia_laws Slovakia Collection of Laws (SLOV-LEX) Research snapshot of official national legislation from SLOV-LEX (Ministry of Justice) Collection of Laws static mirror. Not legal advice. The official gazette / authentic source prevails over this corpus. Snapshot Field Value Snapshot date 2026-09-02 Coverage complete Source SLOV-LEX (Ministry of Justice) Collection of Laws static mirror Collector scrapers/collect_sk.py Laws / instruments 16,799 Articles… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_slovakia_laws.texttext-retrieval100K<n<1M0 likes96 downloads19h agoHugging Face08saillab /alpaca-slovak-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-slovak-cleaned.text10K<n<100K0 likes82 downloads2y agoHugging Face09mbenco /slovak-sft Slovak SFT Dataset A supervised fine-tuning (SFT) dataset for Slovak language instruction following, constructed from two publicly available Slovak resources: saillab/alpaca-slovak-cleaned — Slovak instruction-response pairs TUKE-DeutscheTelekom/skquad — Slovak question answering, rewritten into chat-style prompts Format Each example follows the standard messages format with three turns: { "messages": [ {"role": "system", "content": "Si užitočný slovenský… See the full description on the dataset page: https://huggingface.co/datasets/mbenco/slovak-sft.texttext-generation10K<n<100K0 likes82 downloads5mo agoHugging Face10TUKE-KEMT /slovak-web-qa-pairstext100K<n<1M0 likes77 downloads9mo agoHugging Face11NaiveNeuron /SlovakCOPA Dataset Card for SlovakCOPA Dataset Description SlovakCOPA is a Slovak translation of the COPA (Choices of Plausible Alternatives) dataset, a benchmark for commonsense causal reasoning. COPA is a causal reasoning dataset where given a premise, the model must choose between two alternatives that are either the cause or the effect of the premise. This dataset contains 600 examples (500 test / 100 validation) with parallel translations in English and standard Slovak.… See the full description on the dataset page: https://huggingface.co/datasets/NaiveNeuron/SlovakCOPA.tabularmultiple-choicen<1K1 likes68 downloads5mo agoHugging Face12Plasmoxy /gigatrue-slovak Gigatrue Slovak abstractive summarisation dataset. Synthetic Gigaword dataset translated to Slovak. Same as https://huggingface.co/datasets/Plasmoxy/gigatrue but translated to Slovak using SeamlessM4T-v2 (https://huggingface.co/docs/transformers/en/model_doc/seamless_m4t_v2). Original dataset adapted from https://huggingface.co/datasets/Harvard/gigaword. This work is supported by the EU NextGenerationEU through the Recovery and Resilience Plan for Slovakia under the project No.… See the full description on the dataset page: https://huggingface.co/datasets/Plasmoxy/gigatrue-slovak.textsummarization1M<n<10M2 likes57 downloads2y agoHugging Face13mteb /SlovakSTS SlovakSTS An MTEB dataset Massive Text Embedding Benchmark A professional Slovak translation of the STS Benchmark (STSb), originally part of the GLUE benchmark. The task is Semantic Textual Similarity (STS): given a pair of sentences, the goal is to predict their semantic similarity on a continuous scale from 0 (completely unrelated) to 5 (semantically equivalent). Sentence pairs are drawn from news headlines, image captions, and forum posts. Task category STS… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SlovakSTS.textsentence-similarity1K<n<10K0 likes54 downloads1mo agoHugging Face14NaiveNeuron /slovaksumThe SlovakSum dataset from the SlovakSum: Slovak News Summarization Dataset paper text100K<n<1M8 likes53 downloads3y agoHugging Face15TUKE-KEMT /slovak-triplets Slovak Triplets Dataset This repository contains the Slovak Triplets Dataset, a collection of triplet sentences in the Slovak language designed for training and evaluating document embedding models. Each triplet consists of an anchor sentence, a positive sentence (similar to the anchor), and a negative sentence (dissimilar to the anchor). Data Source The dataset is extracted form the Slovak part of WebFAQ and MQA datasets, which are publicly available collections of… See the full description on the dataset page: https://huggingface.co/datasets/TUKE-KEMT/slovak-triplets.text100K<n<1M1 likes53 downloads8mo agoHugging Face16JZSG /jz_sync_Open-Government-Slovak-filteredtabular100K<n<1M0 likes39 downloads11mo agoHugging Face17Project-AgML /tree_species_classification_slovak_exact_cropped_extended Tree Species Classification Slovak Exact Cropped Extended This dataset provides real RGB images of tree species collected in natural field environments across Slovakia and the Czech Republic. Images were captured using handheld cameras (Sony Alpha 7 and Canon EOS 4000D) during field campaigns in June and August 2022, focusing on bark and foliage for species identification. The dataset contains 1,367 images across 4 classes: European beech, European silver fir, Norway spruce… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/tree_species_classification_slovak_exact_cropped_extended.imageimage-classification1K<n<10K0 likes36 downloads2d agoHugging Face18Adamos3000 /slovak-sft Slovak SFT Dataset A supervised fine-tuning (SFT) dataset for Slovak language instruction following, constructed from two publicly available Slovak resources: saillab/alpaca-slovak-cleaned — Slovak instruction-response pairs TUKE-DeutscheTelekom/skquad — Slovak question answering, rewritten into chat-style prompts Format Each example follows the standard messages format with three turns: { "messages": [ {"role": "system", "content": "Si užitočný slovenský… See the full description on the dataset page: https://huggingface.co/datasets/Adamos3000/slovak-sft.texttext-generation10K<n<100K0 likes32 downloads3mo agoHugging Face19mteb /SlovakSumRetrieval SlovakSumRetrieval An MTEB dataset Massive Text Embedding Benchmark SlovakSum, a Slovak news summarization dataset consisting of over 200 thousand news articles with titles and short abstracts obtained from multiple Slovak newspapers. Originally intended as a summarization task, but since no human annotations were provided here reformulated to a retrieval task. Task category t2t Domains News, Social, Web, Written Reference… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SlovakSumRetrieval.texttext-retrieval1K<n<10K0 likes31 downloads1y agoHugging Face20mteb /SlovakSumURLClustering SlovakSumURLClustering An MTEB dataset Massive Text Embedding Benchmark Clustering of Slovak news articles from SlovakSum dataset based on the URL structure. Articles are organized into 12 editorial categories including sports, culture, economy, health, travel, politics, and technology sections. Task category Clustering (text-to-category) Domains News, Written Reference Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SlovakSumURLClustering.texttext-classification10K<n<100K0 likes30 downloads1mo agoHugging Face21saillab /alpaca_slovak_tacoThis repository contains the dataset used for the TaCo paper. The dataset follows the style outlined in the TaCo paper, as follows: { "instruction": "instruction in xx", "input": "input in xx", "output": "Instruction in English: instruction in en , Response in English: response in en , Response in xx: response in xx " } Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_slovak_taco.text10K<n<100K0 likes29 downloads2y agoHugging Face22shunyalabs /slovak-speech-datasetaudio1K<n<10K0 likes29 downloads1y agoHugging Face23dokato /exam-slovak-mathbioPart of INCLUDE see research: https://arxiv.org/abs/2411.19799 full dataset: https://huggingface.co/datasets/CohereForAI/include-base-44 textmultiple-choicen<1K1 likes28 downloads2y agoHugging Face24simonko912 /slovak-scrapetext1K<n<10K1 likes28 downloads7mo agoHugging Face25slovak-nlp /slovak-pharmacy-drmax-rerankinggated SlovakPharmacyDrMaxReranking An MTEB dataset Massive Text Embedding Benchmark A reranking dataset created from Q&A content collected from DrMax pharmacy website. The dataset consists of questions about medications, health conditions, and pharmaceutical advice, with answers provided by qualified pharmacists. This dataset is designed to evaluate models' ability to rank relevant pharmaceutical information and expert responses. Task category t2t Domains Medical, Web… See the full description on the dataset page: https://huggingface.co/datasets/slovak-nlp/slovak-pharmacy-drmax-reranking.texttext-ranking10K<n<100K0 likes27 downloads4mo agoHugging Face26NLP-FBK /Slovak_Relation_Extractiontexttext-classification1K<n<10K1 likes26 downloads2y agoHugging Face27slovak-nlp /slovak-pharmacy-mojalekaren-rerankinggated SlovakPharmacyMojaLekarenReranking An MTEB dataset Massive Text Embedding Benchmark A reranking dataset created from Q&A content collected from MojaLekaren pharmacy website. The dataset consists of questions about medications, health conditions, and pharmaceutical advice, with answers provided by qualified pharmacists. This dataset is designed to evaluate models' ability to rank relevant pharmaceutical information and expert responses. Task category t2t Domains Medical… See the full description on the dataset page: https://huggingface.co/datasets/slovak-nlp/slovak-pharmacy-mojalekaren-reranking.texttext-ranking1K<n<10K0 likes26 downloads4mo agoHugging Face28DGurgurov /slovak_sa Sentiment Analysis Data for the Slovak Language Dataset Description: This dataset contains a sentiment analysis dataset from Pecar et al. (2019). Data Structure: The data was used for the project on improving word embeddings with graph knowledge for Low Resource Languages. Citation: @inproceedings{pecar-etal-2019-improving, title = "Improving Sentiment Classification in {S}lovak Language", author = "Pecar, Samuel and Simko, Marian and Bielikova, Maria"… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/slovak_sa.texttext-classification1K<n<10K1 likes24 downloads2y agoHugging Face29kiviki /slovaksum-url-clusteringtext10K<n<100K0 likes23 downloads10mo agoHugging Face30mteb /SlovakHateSpeechClassification SlovakHateSpeechClassification An MTEB dataset Massive Text Embedding Benchmark The dataset contains posts from a social network with human annotations for hateful or offensive language in Slovak. Task category t2c Domains Social, Written Reference https://huggingface.co/datasets/TUKE-KEMT/hate_speech_slovak How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SlovakHateSpeechClassification.texttext-classification10K<n<100K0 likes22 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.