CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SlayerLab /polish-dynaword Polish DynaWord A continuously developed, openly-licensed, human-text Polish corpus — a Polish edition in the Dynaword family (Enevoldsen et al., arXiv:2508.02271). v0.2.5 stable · 4,319,200 documents · 9.64B tokens (tiktoken proxy; canonical Llama-3 count at release) · 18 sources Updated: 2026-08-14 v0.3-dev experimental track · quality/diversity workflow, source-gate validation and candidate-data audits. This is development work, not a released corpus version, and it does… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/polish-dynaword.text-generation1M<n<10M23 likes4.7k downloads6d agoHugging Face02agnostic /qwen3-tts-polish-training2 likes4.3k downloads6mo agoHugging Face03Insta360-Research /Matterport3D_polished Matterport3D_polished Matterport3D_Polished is a panoramic dataset derived from Matterport3D, which was introduced in DiT360. This dataset contains 10,000+ high-resolution (2048 x 1024) indoor panoramic images along with corresponding prompts. Compared with the original dataset, it removes the blurred artifacts at both ends, providing clearer and sharper visual details. Which tasks will benefit from our dataset? Text-to-Panorama Generation ⚙️ Getting… See the full description on the dataset page: https://huggingface.co/datasets/Insta360-Research/Matterport3D_polished.image10K<n<100K19 likes1.1k downloads11mo agoHugging Face04datadriven-company /WolneLektury-TTS-Polish WolneLektury-TTS-Polish A large-scale, high-quality Polish speech dataset for text-to-speech and automatic speech recognition. Data Source Derived from Wolne Lektury (Free Readings), a Polish digital library with public domain audiobooks featuring professional voice actors. Dataset Statistics Metric Value Total samples 383,710 Total duration 997 hours Unique narrators 1207 Male samples 294,756 (767h) Female samples 88,945 (230h) Average… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/WolneLektury-TTS-Polish.audiotext-to-speech100K<n<1M2 likes926 downloads8mo agoHugging Face05PleIAs /Polish-PD 🇵🇱 Polish Public Domain 🇵🇱 Polish-Public Domain or Polish-PD is a large collection aiming to aggregate all Polish monographies and periodicals in the public domain. As of March 2024, it is the biggest Polish open corpus. Dataset summary The collection contains 247,491 individual texts making up 2,697,414,811 words recovered from multiple sources, including Internet Archive and various European national libraries and cultural heritage institutions. Each parquet file… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Polish-PD.tabular10K<n<100K5 likes923 downloads3y agoHugging Face06AITextDetect /AI_Polish_cleanFor detecting machine generated text.texttext-classification3 likes375 downloads1y agoHugging Face07hakari-bench /NanoMTEB-Polish NanoMTEB-Polish This dataset is a Nano-style retrieval dataset for HAKARI-bench. NanoMTEB-Polish is a compact Polish retrieval benchmark assembled from Polish MTEB-family retrieval tasks. It includes Polish CQADupStack domains, FiQA, Natural Questions, PUGG information retrieval, and Quora-style duplicate-question retrieval. Usage from datasets import load_dataset dataset_id = "hakari-bench/NanoMTEB-Polish" split = "cqadupstack_android" queries =… See the full description on the dataset page: https://huggingface.co/datasets/hakari-bench/NanoMTEB-Polish.text100K<n<1M0 likes357 downloads3mo agoHugging Face08ptaszynski /PolishCyberbullyingDataset Expert-annotated dataset to study cyberbullying in Polish language This the first publically available expert-annotated dataset containing annotations of cyberbullying and hate-speech in Polish language. Please, read the paper about the dataset for all necessary details. Model The classification model which achieved the highest classification results for the dataset is also released under the following URL. Polbert-CB - Polish BERT trained for Automatic Cyberbullying… See the full description on the dataset page: https://huggingface.co/datasets/ptaszynski/PolishCyberbullyingDataset.tabular10K<n<100K3 likes343 downloads3y agoHugging Face09JohnTdi /polish-llm-sft-pl Polish LLM SFT Dataset PL | Zbiór danych przygotowany z myślą o poprawie i nauczaniu języka polskiego różnych modeli LLM. EN | Dataset prepared to help various LLMs learn and improve their Polish language capabilities. 38 781 sampli / samples · Apache 2.0 · Język / Language: PL (+ pary tłumaczeniowe PL↔EN) Format Każdy sample / every sample: {"messages": [ {"role": "user", "content": "..."}, {"role": "assistant", "content": "..."} ]} Struktura / Structure… See the full description on the dataset page: https://huggingface.co/datasets/JohnTdi/polish-llm-sft-pl.texttext-generation10K<n<100K1 likes257 downloads4mo agoHugging Face10PiotrSty /saos-polish-court-judgments SAOS Speeches Corpus — Orzeczenia Sądów Polskich Korpus orzeczeń sądowych z Systemu Analizy Orzeczeń Sądowych (SAOS), wygenerowany z oficjalnego API SAOS (www.saos.org.pl/api/dump/judgments). Statystyki Metryka Wartość Sądy sądy powszechne, Sąd Najwyższy, Naczelny Sąd Administracyjny, Trybunał Konstytucyjny Typy orzeczeń wyroki, postanowienia, uzasadnienia, zarządzenia, uchwały Rekordy 296,692 Znaki 5,895,441,856 Słowa 855,425,796 Tokeny… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/saos-polish-court-judgments.tabular100K<n<1M0 likes234 downloads2mo agoHugging Face11PiotrSty /impact-psnc-polish-ocr IMPACT-PSNC Polish OCR Diverse Subset Compact, provenance-preserving subset of the Polish IMPACT ground truth released by the Poznan Supercomputing and Networking Center (PSNC). It is intended for OCR experiments on diverse historical Polish printed material. This subset contains: 89 full-page images from 30 source collections; 599 text-region crops derived from PAGE XML polygons; the 89 corresponding original PAGE XML files; page and region transcriptions; document-level train… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/impact-psnc-polish-ocr.imageimage-to-textn<1K0 likes234 downloads8d agoHugging Face12yhsz123 /Matterport3D_polished Matterport3D_polished Matterport3D_Polished is a panoramic dataset derived from Matterport3D, which was introduced in DiT360. This dataset contains 10,000+ high-resolution (2048 x 1024) indoor panoramic images along with corresponding prompts. Compared with the original dataset, it removes the blurred artifacts at both ends, providing clearer and sharper visual details. Which tasks will benefit from our dataset? Text-to-Panorama Generation ⚙️… See the full description on the dataset page: https://huggingface.co/datasets/yhsz123/Matterport3D_polished.image10K<n<100K1 likes226 downloads6mo agoHugging Face13btrkeks /polish-scores Polish Historical-Scan OMR Benchmark A page-level Optical Music Recognition (OMR) evaluation benchmark of 112 real historical score scans, paired with both **kern (Humdrum) and MusicXML ground-truth transcriptions. Derived from the PRAIG/polish-scores dataset, with kern normalization and a manual-fix pass applied. Released as the real-scan half of the Transcoda evaluation suite alongside btrkeks/verovio-synth-omr and the btrkeks/transcoda-59M-zeroshot-v1 model. Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/btrkeks/polish-scores.imageimage-to-textn<1K1 likes186 downloads4mo agoHugging Face14allegro /summarization-polish-summaries-corpustext10K<n<100K5 likes180 downloads5y agoHugging Face15clarin-knext /wsd_polish_datasetsPolish WSD training data manually annotated by experts according to plWordNet-4.2.token-classification1M<n<10M0 likes172 downloads3y agoHugging Face16SlayerLab /fabryka-track-polish-mix Fabryka Track Polish training mix (100 MB/source pack) This dataset is the verified corpus pack used by track.fabryka.ai for training-pipeline tests. It contains UTF-8 text samples plus one JSON metadata file per source and catalog.json. The bounded sources were materialized from fixed Hugging Face revisions. bytes in the catalog is the exact UTF-8 byte size; the current Track byte-token trainer counts one byte as one training token. This is a reproducible workflow pack, not a… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/fabryka-track-polish-mix.tabularn<1K1 likes143 downloads13d agoHugging Face17allegro /polish-question-passage-pairstext10K<n<100K5 likes136 downloads5y agoHugging Face18Paul /hatecheck-polish Dataset Card for Multilingual HateCheck Dataset Description Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish. For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate. This allows for targeted diagnostic insights into model performance. For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-polish.tabulartext-classification1K<n<10K3 likes136 downloads4y agoHugging Face19PiotrSty /wolne-lektury-polish-literature-corpus Wolne Lektury Polish Literature Corpus Dataset Description A comprehensive corpus of Polish literary works from Wolne Lektury — a free digital library of public domain literature. All texts are in the public domain. The corpus was collected via the official REST API (https://wolnelektury.pl/api/), including full text of each work, metadata (author, epoch, genre, kind), and language information. Statistics Metric Value Records 7,316… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/wolne-lektury-polish-literature-corpus.tabulartext-generation10K<n<100K0 likes117 downloads3mo agoHugging Face20SlayerLab /polish-dynaword-mix-extended-500M Polish DynaWord Mix — Extended (~500M tokens) Maintained by Arkadiusz Słota · SlayerLab A curated, openly-licensed Polish text corpus for language-model pretraining and research baselines. Built on top of the open polish-dynaword lineage and extended to ~500 million tokens (32k BPE) with additional curated, license-compatible sources and a documented cleaning + PII-scrubbing pipeline. Why this exists: most large Polish web corpora are legal/parliamentary-heavy and carry… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/polish-dynaword-mix-extended-500M.texttext-generation10M<n<100M0 likes114 downloads1mo agoHugging Face21cpral /Polski_Polish_SFT_poziomka-fun-rp-v11 poziomka-fun-rp-v11 Polski zbiór SFT oparty na poziomka-fun-rp-v10, z przetasowanymi danymi i opcjonalnym sterowaniem reasoning przez prefiks odpowiedzi asystenta, wzorem Qwen. Cały reasoning z v10 pozostaje w wiadomościach i w tekście treningowym. Nie wyłączano losowo rozumowania, nie przenoszono go do archiwum ani nie dopisywano komend do użytkownika. System prompty pozostają bez zmian. Rozmiar i źródła Podział Rozmowy Format v10 Wybrane do formatu… See the full description on the dataset page: https://huggingface.co/datasets/cpral/Polski_Polish_SFT_poziomka-fun-rp-v11.text1M<n<10M0 likes100 downloads13d agoHugging Face22PRAIG /polish-scoresimagen<1K0 likes98 downloads11mo agoHugging Face23JohnTdi /bielik-distill-polish-10k bielik-distill-polish-10k Polish instruction-tuning dataset with 10,304 samples generated via response-level knowledge distillation from Bielik-11B-v3.0-Instruct (SpeakLeash, Apache 2.0). Covers Polish history, culture, politics, science, geography, idioms, and general reasoning. Multi-pass quality control: factual corrections, topic filtering (Poland/Europe focus), truncation removal (~9% of raw data removed). Format { "messages": [ {"role": "user"… See the full description on the dataset page: https://huggingface.co/datasets/JohnTdi/bielik-distill-polish-10k.texttext-generation10K<n<100K1 likes92 downloads4mo agoHugging Face24NASK-PIB /Reassess-Polish-Medical-Examstext10K<n<100K0 likes89 downloads14d agoHugging Face25s512757 /polish-tedx-asr-eval Polish-TEDx-ASR-Eval A dataset for evaluating automatic speech recognition (ASR) systems for Polish in the domain of TEDx public talks. Contains audio segments from Polish TEDx talks available on YouTube (CC BY-NC-ND 4.0) and synthetic speech generated with KugelAudio (MIT), with manually created and cross-verified transcriptions. Created as part of the course "Workshops on Evaluation of Speech Recognition Systems" (ZWESUI, AMU 2026) by Group 1. Statistics… See the full description on the dataset page: https://huggingface.co/datasets/s512757/polish-tedx-asr-eval.audioautomatic-speech-recognitionn<1K0 likes88 downloads3mo agoHugging Face26open-llm-leaderboard-old /details_Yuma42__KangalKhan-PolishedRuby-7B Dataset Card for Evaluation run of Yuma42/KangalKhan-PolishedRuby-7B Dataset automatically created during the evaluation run of model Yuma42/KangalKhan-PolishedRuby-7B on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Yuma42__KangalKhan-PolishedRuby-7B.0 likes82 downloads2y agoHugging Face27shunyalabs /polish-speech-datasetaudio1K<n<10K0 likes81 downloads1y agoHugging Face28DebasishDhal99 /german-polish-paired-placenames Dataset Summary This dataset contains the German and Polish names for almost 10k places in Poland. It has been generated using this code. Many of these names are related to each other. Some German names are literal translation of the Polish names, some are phonetic modifications while some are unrelated. Dataset Creation Source Data German wiki page texttranslation1K<n<10K0 likes72 downloads3y agoHugging Face29directtt /polish_presidential_debate Polish Presidential ASR Dataset Public domain (2025) Source: DEBATA PREZYDENCKA TVP | 12.05.2025 Description Studio-recorded utterances by 13 Polish presidential debate candidates, each providing 15 audio samples (read or spontaneous). Audio is in 16 kHz WAV format. The dataset is designed for automatic speech recognition (ASR) tasks, particularly in the context of Polish language processing in political domain. File Structure ├── README.md ├── audio/ │ └──… See the full description on the dataset page: https://huggingface.co/datasets/directtt/polish_presidential_debate.automatic-speech-recognitionn<1K2 likes72 downloads1y agoHugging Face30PiotrSty /ted-polish-procurement-notices TED Polish procurement notices Polish narrative text rebuilt from official TED procurement-notice records. Indexed notices: 20,588 Search API records: 20,588 XML fallbacks: 1,255 Retained documents: 15,544 Tokens: 42,475,098 (cl100k_base proxy) Dates: 2023-02-15 to 2024-05-15 Contracting-authority attribution: 100.0% The pinned PleIAs mirror is an identifier index only because its Polish preview contains replacement-character encoding damage. Released text comes from official… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/ted-polish-procurement-notices.texttext-generation10K<n<100K0 likes72 downloads14d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.