datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
legendy-i-padanni_original
Легенды і паданні — арыгінальнае аўдыё
Мова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
legendy-i-padanni
Доўгасць аўдыё
1h49m
Радкоў у датасеце
533
Структура
Кожны радок змяшчае:
audio — арыгінальны аўдыёзапіс
text — транскрыпцыя
chunk_uid — унікальны ідэнтыфікатар… See the full description on the dataset page: https://huggingface.co/datasets/fosters/legendy-i-padanni_original.Leonardo_Legends.VoiceLineslegendy-i-padanni
Легенды і паданні
Мова / Language: Беларуская (Belarusian)
Частка калекцыі Ministerskija —
выраўнаваныя аўдыёзапісы беларускіх аўдыёкніг з транскрыпцыямі.
Апублікаваных радкоў (HF)
533
Доўгасць аўдыё
1h49m
Парог даверу
≥ 0.95
Структура
Кожны радок змяшчае:
audio — аўдыёфрагмент (~15 с)
text — транскрыпцыя (Gemini + ASR выраўнаванне)
chunk_uid — унікальны ідэнтыфікатар фрагмента
Апрацоўка
Аўдыёкніга разбіта на кароткія… See the full description on the dataset page: https://huggingface.co/datasets/fosters/legendy-i-padanni.legendy-i-padanni_all
Легенды і паданні
Мова / Language: Беларуская (Belarusian)
Аўдыё нарэзана з арыгінальнага запісу ў зыходнай частаце дыскрэтызацыі (native), мона, фрагменты да 30 секунд.
Частка калекцыі Belarusian Audiobooks (native).
Радкоў у датасеце
533
Працягласць
1 гадз 41 хв
Частата дыскрэтызацыі
44100 Hz
Каналы
мона
Даўжыня фрагмента
да 30 с
Структура
Кожны радок змяшчае:
audio — аўдыёфрагмент (native SR, мона, ≤30 с)
text — транскрыпцыя (Gemini +… See the full description on the dataset page: https://huggingface.co/datasets/fosters/legendy-i-padanni_all.legendy-i-padanni-viktar-manaeu
Легенды і паданні
Metadata
Author:
Title: Легенды і паданні
Narrator: Віктар Манаеў
Source Group: Дзіцячыя
Source: https://www.twirpx.com/file/1172032/
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders.… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/legendy-i-padanni-viktar-manaeu.uladzimir-karatkevich-sivaia-legenda-andrei-kaliada
Сівая легенда
Metadata
Author: Уладзімір Караткевіч
Title: Сівая легенда
Narrator: Андрэй Каляда
Source Group: Аўдыёкнігі
Source: http://rutracker.org/forum/viewtopic.php?t=3678922
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/uladzimir-karatkevich-sivaia-legenda-andrei-kaliada.legendy-i-padanni-input
Легенды і паданні
This is a Hugging Face Parquet input dataset for an audio pipeline.
Source dataset: archivartaunik/legendy-i-padanni
Format
config: default
split: train
format: parquet
id column: id
audio column: audio
rows: 42
shards: 1
language: be
The audio column is embedded into Parquet as Hugging Face Audio:
audio = {
"path": "file.mp3",
"bytes": b"..."
}
Columns
id
audio
title
language
file_name
filename
Author
Title
Narrator… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/legendy-i-padanni-input.legendy-i-padanni
Легенды і паданні
Metadata
Author:
Title: Легенды і паданні
Narrator:
Source Group: Дзіцячыя
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders.
Target maximum split size: about 250 MB.
Each split… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/legendy-i-padanni.
