polish
Datasets
All datasets matching “polish”polish-dynaword
Polish DynaWord
A continuously developed, openly-licensed, human-text Polish corpus — a Polish
edition in the Dynaword
family (Enevoldsen et al., arXiv:2508.02271).
v0.2.5 stable · 4,319,200 documents · 9.64B tokens
(tiktoken proxy; canonical Llama-3 count at release) · 18 sources
Updated: 2026-08-14
v0.3-dev experimental track · quality/diversity workflow, source-gate
validation and candidate-data audits. This is development work, not a released
corpus version, and it does… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/polish-dynaword.qwen3-tts-polish-trainingMatterport3D_polished
Matterport3D_polished
Matterport3D_Polished is a panoramic dataset derived from Matterport3D, which was introduced in DiT360.
This dataset contains 10,000+ high-resolution (2048 x 1024) indoor panoramic images along with corresponding prompts.
Compared with the original dataset, it removes the blurred artifacts at both ends, providing clearer and sharper visual details.
Which tasks will benefit from our dataset?
Text-to-Panorama Generation
⚙️ Getting… See the full description on the dataset page: https://huggingface.co/datasets/Insta360-Research/Matterport3D_polished.WolneLektury-TTS-Polish
WolneLektury-TTS-Polish
A large-scale, high-quality Polish speech dataset for text-to-speech and automatic speech recognition.
Data Source
Derived from Wolne Lektury (Free Readings), a Polish digital library with public domain audiobooks featuring professional voice actors.
Dataset Statistics
Metric
Value
Total samples
383,710
Total duration
997 hours
Unique narrators
1207
Male samples
294,756 (767h)
Female samples
88,945 (230h)
Average… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/WolneLektury-TTS-Polish.Polish-PD
🇵🇱 Polish Public Domain 🇵🇱
Polish-Public Domain or Polish-PD is a large collection aiming to aggregate all Polish monographies and periodicals in the public domain. As of March 2024, it is the biggest Polish open corpus.
Dataset summary
The collection contains 247,491 individual texts making up 2,697,414,811 words recovered from multiple sources, including Internet Archive and various European national libraries and cultural heritage institutions. Each parquet file… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Polish-PD.AI_Polish_cleanFor detecting machine generated text.
