datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
knihi-be-arlou_sny_impieratara_all
AudioSet Pipeline Output
Мова / Language: Беларуская (Belarusian)
Аўдыё нарэзана з арыгінальнага запісу ў зыходнай частаце дыскрэтызацыі (native), мона, фрагменты да 30 секунд.
Частка калекцыі Belarusian Audiobooks (native).
Радкоў у датасеце
870
Працягласць
2 гадз 50 хв
Частата дыскрэтызацыі
44100 Hz
Каналы
мона
Даўжыня фрагмента
да 30 с
Структура
Кожны радок змяшчае:
audio — аўдыёфрагмент (native SR, мона, ≤30 с)
text — транскрыпцыя… See the full description on the dataset page: https://huggingface.co/datasets/fosters/knihi-be-arlou_sny_impieratara_all.knihi-be-arlou_sny_impieratara_output_original
AudioSet Pipeline Output — арыгінальнае аўдыё
Мова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
knihi-be-arlou_sny_impieratara_output
Структура
Кожны радок змяшчае:
audio — арыгінальны аўдыёзапіс
text — транскрыпцыя
chunk_uid — унікальны ідэнтыфікатар
Ліцэнзія / License… See the full description on the dataset page: https://huggingface.co/datasets/fosters/knihi-be-arlou_sny_impieratara_output_original.UGAkan-ImpairedSpeech
UGAkan-ImpairedSpeechData (Research Working Copy)
Personal working copy for ASR research.
Original dataset: UGAkan-ImpairedSpeechData (University of Ghana).
Structure on this repo
Metadata Columns
Column
Description
filename
Audio filename
aetiology
Raw aetiology label
aetiology_key
Normalised key
gender
Speaker gender
speaker_id
Unique speaker ID
environment
Recording environment
transcription
Akan transcription
duration… See the full description on the dataset page: https://huggingface.co/datasets/Kennethdot/UGAkan-ImpairedSpeech.
