datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PolishCyberbullyingDataset
Expert-annotated dataset to study cyberbullying in Polish language
This the first publically available expert-annotated dataset containing annotations of cyberbullying and hate-speech in Polish language.
Please, read the paper about the dataset for all necessary details.
Model
The classification model which achieved the highest classification results for the dataset is also released under the following URL.
Polbert-CB - Polish BERT trained for Automatic Cyberbullying… See the full description on the dataset page: https://huggingface.co/datasets/ptaszynski/PolishCyberbullyingDataset.summarization-polish-summaries-corpuspolish-question-passage-pairshatecheck-polish
Dataset Card for Multilingual HateCheck
Dataset Description
Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish.
For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate.
This allows for targeted diagnostic insights into model performance.
For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-polish.polish-tedx-asr-eval
Polish-TEDx-ASR-Eval
A dataset for evaluating automatic speech recognition (ASR) systems for Polish in the domain of TEDx public talks.
Contains audio segments from Polish TEDx talks available on YouTube (CC BY-NC-ND 4.0) and synthetic speech generated with KugelAudio (MIT), with manually created and cross-verified transcriptions. Created as part of the course "Workshops on Evaluation of Speech Recognition Systems" (ZWESUI, AMU 2026) by Group 1.
Statistics… See the full description on the dataset page: https://huggingface.co/datasets/s512757/polish-tedx-asr-eval.german-polish-paired-placenames
Dataset Summary
This dataset contains the German and Polish names for almost 10k places in Poland. It has been generated using this code.
Many of these names are related to each other. Some German names are literal translation of the Polish names, some are phonetic modifications while some are unrelated.
Dataset Creation
Source Data
German wiki page
polish-newsThis dataset contains more than 250k articles obtained from polish news site tvp.info.pl.
Main purpouse of collecting the data was to create a transformer-based model for text summarization.
Columns:
link - link to article
title - original title of the article
headline - lead/headline of the article - first paragraph of the article visible directly from the page
content - full textual contents of the article
Link to original repo: https://github.com/WiktorSob/scraper-tvp
Download the data:… See the full description on the dataset page: https://huggingface.co/datasets/WiktorS/polish-news.polish-qa-generalPolish_sentimentABCD-polish-QA-with-CoTPolish-Speech-Dataset
🎧 Polish Speech Dataset
The Polish Speech Dataset is a high-quality speech audio dataset designed to support advanced AI and machine learning systems with structured and diverse audio data. It includes 121 hours of recorded speech data across 688 files, provided in MP3 and WAV formats, with a total size of 207 MB. This carefully curated audio dataset ensures balanced and representative voice data, with 51% female and 49% male speakers, and age distribution spanning from 18 to 50+… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Polish-Speech-Dataset.allegro-polish-summaries-corpus-llama2-2000rowswikipedia-polish-qaIMPORTANT:
i made this some time ago, so i dont feel like fixing it, but there is a massive bias in correct anwsers (about 93% of corret anwsers are either B or C). that is relativly easy to fix, but if not noticed, may fuck up your model.
polish-sentiment-datasetpolish_sentencespolish-presidential-debate
