datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
polish-dynaword
Polish DynaWord
A continuously developed, openly-licensed, human-text Polish corpus — a Polish
edition in the Dynaword
family (Enevoldsen et al., arXiv:2508.02271).
v0.2.5 stable · 4,319,200 documents · 9.64B tokens
(tiktoken proxy; canonical Llama-3 count at release) · 18 sources
Updated: 2026-08-14
v0.3-dev experimental track · quality/diversity workflow, source-gate
validation and candidate-data audits. This is development work, not a released
corpus version, and it does… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/polish-dynaword.qwen3-tts-polish-trainingMatterport3D_polished
Matterport3D_polished
Matterport3D_Polished is a panoramic dataset derived from Matterport3D, which was introduced in DiT360.
This dataset contains 10,000+ high-resolution (2048 x 1024) indoor panoramic images along with corresponding prompts.
Compared with the original dataset, it removes the blurred artifacts at both ends, providing clearer and sharper visual details.
Which tasks will benefit from our dataset?
Text-to-Panorama Generation
⚙️ Getting… See the full description on the dataset page: https://huggingface.co/datasets/Insta360-Research/Matterport3D_polished.WolneLektury-TTS-Polish
WolneLektury-TTS-Polish
A large-scale, high-quality Polish speech dataset for text-to-speech and automatic speech recognition.
Data Source
Derived from Wolne Lektury (Free Readings), a Polish digital library with public domain audiobooks featuring professional voice actors.
Dataset Statistics
Metric
Value
Total samples
383,710
Total duration
997 hours
Unique narrators
1207
Male samples
294,756 (767h)
Female samples
88,945 (230h)
Average… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/WolneLektury-TTS-Polish.Polish-PD
🇵🇱 Polish Public Domain 🇵🇱
Polish-Public Domain or Polish-PD is a large collection aiming to aggregate all Polish monographies and periodicals in the public domain. As of March 2024, it is the biggest Polish open corpus.
Dataset summary
The collection contains 247,491 individual texts making up 2,697,414,811 words recovered from multiple sources, including Internet Archive and various European national libraries and cultural heritage institutions. Each parquet file… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Polish-PD.AI_Polish_cleanFor detecting machine generated text.NanoMTEB-Polish
NanoMTEB-Polish
This dataset is a Nano-style retrieval dataset for HAKARI-bench.
NanoMTEB-Polish is a compact Polish retrieval benchmark assembled from Polish MTEB-family retrieval tasks. It includes Polish CQADupStack domains, FiQA, Natural Questions, PUGG information retrieval, and Quora-style duplicate-question retrieval.
Usage
from datasets import load_dataset
dataset_id = "hakari-bench/NanoMTEB-Polish"
split = "cqadupstack_android"
queries =… See the full description on the dataset page: https://huggingface.co/datasets/hakari-bench/NanoMTEB-Polish.PolishCyberbullyingDataset
Expert-annotated dataset to study cyberbullying in Polish language
This the first publically available expert-annotated dataset containing annotations of cyberbullying and hate-speech in Polish language.
Please, read the paper about the dataset for all necessary details.
Model
The classification model which achieved the highest classification results for the dataset is also released under the following URL.
Polbert-CB - Polish BERT trained for Automatic Cyberbullying… See the full description on the dataset page: https://huggingface.co/datasets/ptaszynski/PolishCyberbullyingDataset.polish-llm-sft-pl
Polish LLM SFT Dataset
PL | Zbiór danych przygotowany z myślą o poprawie i nauczaniu języka polskiego różnych modeli LLM.
EN | Dataset prepared to help various LLMs learn and improve their Polish language capabilities.
38 781 sampli / samples · Apache 2.0 · Język / Language: PL (+ pary tłumaczeniowe PL↔EN)
Format
Każdy sample / every sample:
{"messages": [
{"role": "user", "content": "..."},
{"role": "assistant", "content": "..."}
]}
Struktura / Structure… See the full description on the dataset page: https://huggingface.co/datasets/JohnTdi/polish-llm-sft-pl.saos-polish-court-judgments
SAOS Speeches Corpus — Orzeczenia Sądów Polskich
Korpus orzeczeń sądowych z Systemu Analizy Orzeczeń Sądowych (SAOS), wygenerowany z oficjalnego API SAOS (www.saos.org.pl/api/dump/judgments).
Statystyki
Metryka
Wartość
Sądy
sądy powszechne, Sąd Najwyższy, Naczelny Sąd Administracyjny, Trybunał Konstytucyjny
Typy orzeczeń
wyroki, postanowienia, uzasadnienia, zarządzenia, uchwały
Rekordy
296,692
Znaki
5,895,441,856
Słowa
855,425,796
Tokeny… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/saos-polish-court-judgments.impact-psnc-polish-ocr
IMPACT-PSNC Polish OCR Diverse Subset
Compact, provenance-preserving subset of the Polish IMPACT ground truth released by the Poznan Supercomputing and Networking Center (PSNC). It is intended for OCR experiments on diverse historical Polish printed material.
This subset contains:
89 full-page images from 30 source collections;
599 text-region crops derived from PAGE XML polygons;
the 89 corresponding original PAGE XML files;
page and region transcriptions;
document-level train… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/impact-psnc-polish-ocr.Matterport3D_polished
Matterport3D_polished
Matterport3D_Polished is a panoramic dataset derived from Matterport3D, which was introduced in DiT360.
This dataset contains 10,000+ high-resolution (2048 x 1024) indoor panoramic images along with corresponding prompts.
Compared with the original dataset, it removes the blurred artifacts at both ends, providing clearer and sharper visual details.
Which tasks will benefit from our dataset?
Text-to-Panorama Generation
⚙️… See the full description on the dataset page: https://huggingface.co/datasets/yhsz123/Matterport3D_polished.polish-scores
Polish Historical-Scan OMR Benchmark
A page-level Optical Music Recognition (OMR) evaluation benchmark of 112 real
historical score scans, paired with both **kern (Humdrum) and MusicXML
ground-truth transcriptions.
Derived from the PRAIG/polish-scores
dataset, with kern normalization and a manual-fix pass applied. Released as
the real-scan half of the Transcoda evaluation suite alongside
btrkeks/verovio-synth-omr
and the btrkeks/transcoda-59M-zeroshot-v1
model.
Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/btrkeks/polish-scores.summarization-polish-summaries-corpuswsd_polish_datasetsPolish WSD training data manually annotated by experts according to plWordNet-4.2.fabryka-track-polish-mix
Fabryka Track Polish training mix (100 MB/source pack)
This dataset is the verified corpus pack used by track.fabryka.ai for training-pipeline tests.
It contains UTF-8 text samples plus one JSON metadata file per source and catalog.json.
The bounded sources were materialized from fixed Hugging Face revisions. bytes in the catalog is the exact UTF-8 byte size; the current Track byte-token trainer counts one byte as one training token.
This is a reproducible workflow pack, not a… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/fabryka-track-polish-mix.polish-question-passage-pairshatecheck-polish
Dataset Card for Multilingual HateCheck
Dataset Description
Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish.
For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate.
This allows for targeted diagnostic insights into model performance.
For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-polish.wolne-lektury-polish-literature-corpus
Wolne Lektury Polish Literature Corpus
Dataset Description
A comprehensive corpus of Polish literary works from Wolne Lektury — a free digital library of public domain literature. All texts are in the public domain.
The corpus was collected via the official REST API (https://wolnelektury.pl/api/), including full text of each work, metadata (author, epoch, genre, kind), and language information.
Statistics
Metric
Value
Records
7,316… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/wolne-lektury-polish-literature-corpus.polish-dynaword-mix-extended-500M
Polish DynaWord Mix — Extended (~500M tokens)
Maintained by Arkadiusz Słota · SlayerLab
A curated, openly-licensed Polish text corpus for language-model pretraining and
research baselines. Built on top of the open polish-dynaword lineage and extended to
~500 million tokens (32k BPE) with additional curated, license-compatible sources
and a documented cleaning + PII-scrubbing pipeline.
Why this exists: most large Polish web corpora are legal/parliamentary-heavy and
carry… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/polish-dynaword-mix-extended-500M.Polski_Polish_SFT_poziomka-fun-rp-v11
poziomka-fun-rp-v11
Polski zbiór SFT oparty na poziomka-fun-rp-v10, z przetasowanymi danymi i opcjonalnym sterowaniem reasoning przez prefiks odpowiedzi asystenta, wzorem Qwen.
Cały reasoning z v10 pozostaje w wiadomościach i w tekście treningowym. Nie wyłączano losowo rozumowania, nie przenoszono go do archiwum ani nie dopisywano komend do użytkownika. System prompty pozostają bez zmian.
Rozmiar i źródła
Podział
Rozmowy
Format v10
Wybrane do formatu… See the full description on the dataset page: https://huggingface.co/datasets/cpral/Polski_Polish_SFT_poziomka-fun-rp-v11.polish-scoresbielik-distill-polish-10k
bielik-distill-polish-10k
Polish instruction-tuning dataset with 10,304 samples generated via response-level knowledge distillation from Bielik-11B-v3.0-Instruct (SpeakLeash, Apache 2.0).
Covers Polish history, culture, politics, science, geography, idioms, and general reasoning. Multi-pass quality control: factual corrections, topic filtering (Poland/Europe focus), truncation removal (~9% of raw data removed).
Format
{
"messages": [
{"role": "user"… See the full description on the dataset page: https://huggingface.co/datasets/JohnTdi/bielik-distill-polish-10k.Reassess-Polish-Medical-Examspolish-tedx-asr-eval
Polish-TEDx-ASR-Eval
A dataset for evaluating automatic speech recognition (ASR) systems for Polish in the domain of TEDx public talks.
Contains audio segments from Polish TEDx talks available on YouTube (CC BY-NC-ND 4.0) and synthetic speech generated with KugelAudio (MIT), with manually created and cross-verified transcriptions. Created as part of the course "Workshops on Evaluation of Speech Recognition Systems" (ZWESUI, AMU 2026) by Group 1.
Statistics… See the full description on the dataset page: https://huggingface.co/datasets/s512757/polish-tedx-asr-eval.details_Yuma42__KangalKhan-PolishedRuby-7B
Dataset Card for Evaluation run of Yuma42/KangalKhan-PolishedRuby-7B
Dataset automatically created during the evaluation run of model Yuma42/KangalKhan-PolishedRuby-7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Yuma42__KangalKhan-PolishedRuby-7B.polish-speech-datasetgerman-polish-paired-placenames
Dataset Summary
This dataset contains the German and Polish names for almost 10k places in Poland. It has been generated using this code.
Many of these names are related to each other. Some German names are literal translation of the Polish names, some are phonetic modifications while some are unrelated.
Dataset Creation
Source Data
German wiki page
polish_presidential_debate
Polish Presidential ASR Dataset
Public domain (2025)
Source: DEBATA PREZYDENCKA TVP | 12.05.2025
Description
Studio-recorded utterances by 13 Polish presidential debate candidates, each providing 15 audio samples (read or spontaneous). Audio is in 16 kHz WAV format.
The dataset is designed for automatic speech recognition (ASR) tasks, particularly in the context of Polish language processing in political domain.
File Structure
├── README.md
├── audio/
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/directtt/polish_presidential_debate.ted-polish-procurement-notices
TED Polish procurement notices
Polish narrative text rebuilt from official TED procurement-notice records.
Indexed notices: 20,588
Search API records: 20,588
XML fallbacks: 1,255
Retained documents: 15,544
Tokens: 42,475,098 (cl100k_base proxy)
Dates: 2023-02-15 to 2024-05-15
Contracting-authority attribution: 100.0%
The pinned PleIAs mirror is an identifier index only because its Polish preview contains replacement-character encoding damage. Released text comes from official… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/ted-polish-procurement-notices.
