datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WolneLektury-TTS-Polish
WolneLektury-TTS-Polish
A large-scale, high-quality Polish speech dataset for text-to-speech and automatic speech recognition.
Data Source
Derived from Wolne Lektury (Free Readings), a Polish digital library with public domain audiobooks featuring professional voice actors.
Dataset Statistics
Metric
Value
Total samples
383,710
Total duration
997 hours
Unique narrators
1207
Male samples
294,756 (767h)
Female samples
88,945 (230h)
Average… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/WolneLektury-TTS-Polish.polish-tedx-asr-eval
Polish-TEDx-ASR-Eval
A dataset for evaluating automatic speech recognition (ASR) systems for Polish in the domain of TEDx public talks.
Contains audio segments from Polish TEDx talks available on YouTube (CC BY-NC-ND 4.0) and synthetic speech generated with KugelAudio (MIT), with manually created and cross-verified transcriptions. Created as part of the course "Workshops on Evaluation of Speech Recognition Systems" (ZWESUI, AMU 2026) by Group 1.
Statistics… See the full description on the dataset page: https://huggingface.co/datasets/s512757/polish-tedx-asr-eval.Polish-Speech-Dataset
🎧 Polish Speech Dataset
The Polish Speech Dataset is a high-quality speech audio dataset designed to support advanced AI and machine learning systems with structured and diverse audio data. It includes 121 hours of recorded speech data across 688 files, provided in MP3 and WAV formats, with a total size of 207 MB. This carefully curated audio dataset ensures balanced and representative voice data, with 51% female and 49% male speakers, and age distribution spanning from 18 to 50+… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Polish-Speech-Dataset.YodaLingua-Polish
YodaLingua-Polish
YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Polish portion of the multilingual YodaLingua collection.
🧾 Dataset Overview
Property
Value
Total clips
329,740 audio–transcription pairs
Total duration
893 hours
Speakers
11,357 distinct speakers
Audio format
MP3 • mono • 24 kHz • 16-bit… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Polish.
