datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
astryd_lindgren_braty_lvinae_sertsa_all
Браты Ільвінае сэрца
Аўтар / Author: Астрыд ЛіндгрэнМова / Language: Беларуская (Belarusian)
Аўдыё нарэзана з арыгінальнага запісу ў зыходнай частаце дыскрэтызацыі (native), мона, фрагменты да 30 секунд.
Частка калекцыі Belarusian Audiobooks (native).
Радкоў у датасеце
2,116
Працягласць
6 гадз 31 хв
Частата дыскрэтызацыі
44100 Hz
Каналы
мона
Даўжыня фрагмента
да 30 с
Структура
Кожны радок змяшчае:
audio — аўдыёфрагмент (native SR, мона… See the full description on the dataset page: https://huggingface.co/datasets/fosters/astryd_lindgren_braty_lvinae_sertsa_all.astryd_lindgren_braty_lvinae_sertsa_output_original
Браты Ільвінае сэрца — арыгінальнае аўдыё
Аўтар / Author: Астрыд ЛіндгрэнМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
astryd_lindgren_braty_lvinae_sertsa_output
Доўгасць аўдыё
7h06m
Радкоў у датасеце
1,835
Структура
Кожны радок змяшчае:
audio — арыгінальны… See the full description on the dataset page: https://huggingface.co/datasets/fosters/astryd_lindgren_braty_lvinae_sertsa_output_original.asturian-speech-dataset
Dataset Card
Dataset Description
This dataset contains parts of data from Google FLEURS (Few-shot Learning Evaluation of Universal Representations of Speech).
Dataset Summary
The dataset includes audio recordings sampled at 16kHz along with their corresponding transcripts. It is split into training, validation, and test sets for speech recognition and related tasks.
Dataset Structure
Data Fields
audio: An audio file with a sampling rate… See the full description on the dataset page: https://huggingface.co/datasets/shunyalabs/asturian-speech-dataset.
