datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
common-voice-pl-text
Common Voice Polish validated text v26.0
Versioned text-only research snapshot prepared for Polish DynaWord. It contains
45,043 unique Polish sentences associated with validated Common Voice
recordings and 994,922 cl100k_base proxy tokens.
The dataset is derived from Mozilla Common Voice Scripted Speech 26.0
through the pinned mirror Peacockery/common-voice-scripted-speech-26@b4d8b94d43831475de59a455345acf6945cfd66e. The source is
distributed under CC0-1.0.
Files… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/common-voice-pl-text.mozilla-common-voice-23-bel-texts-export
