danijelkorzinek/ClarinStudioPL
CLARIN-PL Polish Studio Corpus The corpus was created somewhere in 2014-2015 by recording a group of few hundred volunteer speakers reading a few dozen sentences each. The total size of the corpus is ~56 hours. Due to the manner of recording, the transcription accuracy is very high, but the manner of speech is not spontaneous. This corpus is best compared to something like TIMIT, possibly CommonVoice. It is different from CommonVoice in that it is recorded in a controlled… See the full description on the dataset page: https://huggingface.co/datasets/danijelkorzinek/ClarinStudioPL.
2130
../
dev-00000-of-00002.parquetdownload
dev-00001-of-00002.parquetdownload
test-00000-of-00002.parquetdownload
test-00001-of-00002.parquetdownload
train-00000-of-00011.parquetdownload
train-00001-of-00011.parquetdownload
train-00002-of-00011.parquetdownload
train-00003-of-00011.parquetdownload
train-00004-of-00011.parquetdownload
train-00005-of-00011.parquetdownload
train-00006-of-00011.parquetdownload
train-00007-of-00011.parquetdownload
train-00008-of-00011.parquetdownload
train-00009-of-00011.parquetdownload
train-00010-of-00011.parquetdownload
