datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Roberto_Carlosdzintra_shultse_robertsik_output_original
Робэртсік — арыгінальнае аўдыё
Аўтар / Author: Дзінтра ШультсэМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
dzintra_shultse_robertsik_output
Доўгасць аўдыё
0h22m
Радкоў у датасеце
98
Структура
Кожны радок змяшчае:
audio — арыгінальны аўдыёзапіс
text — транскрыпцыя… See the full description on the dataset page: https://huggingface.co/datasets/fosters/dzintra_shultse_robertsik_output_original.Vozesus-robert-higgs-metadata4-v2dzintra_shultse_robertsik_all
Робэртсік
Аўтар / Author: Дзінтра ШультсэМова / Language: Беларуская (Belarusian)
Аўдыё нарэзана з арыгінальнага запісу ў зыходнай частаце дыскрэтызацыі (native), мона, фрагменты да 30 секунд.
Частка калекцыі Belarusian Audiobooks (native).
Радкоў у датасеце
112
Працягласць
22 хв
Частата дыскрэтызацыі
44100 Hz
Каналы
мона
Даўжыня фрагмента
да 30 с
Структура
Кожны радок змяшчае:
audio — аўдыёфрагмент (native SR, мона, ≤30 с)
text —… See the full description on the dataset page: https://huggingface.co/datasets/fosters/dzintra_shultse_robertsik_all.us-robert-higgs-metadata4-v1scotus-john_g_roberts_jr-audio
SCOTUS-sim audio: john_g_roberts_jr
Per-utterance audio clips from Oyez oral-argument mp3s, sliced at
the start_time / stop_time timestamps stored in the companion
scotus-sim/scotus-john_g_roberts_jr-training dataset.
Alignment
clip_NNNNN.wav in the tarball corresponds exactly to
audio_segments.jsonl[NNNNN] in the training companion dataset.
In metadata.jsonl each row carries the same 0-padded index in idx.
This supersedes the v1 tarball, which had systematic… See the full description on the dataset page: https://huggingface.co/datasets/scotus-sim/scotus-john_g_roberts_jr-audio.robertodataset2roberto90datasetRobertoCarlos9091minhavozariyanvoiceresult_with_finetuned_taggenv2_10epoch_encoder_embeddings_decoder_roberta
Dataset Card for "result_with_finetuned_taggenv2_10epoch_encoder_embeddings_decoder_roberta"
More Information needed
Robertorobert-iang-ssekchy-dreva
Сьсекчы дрэва
Metadata
Author: Робэрт Янг
Title: Сьсекчы дрэва
Narrator:
Source Group: Аўдыёкнігі
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders.
Target maximum split size: about 250 MB.
Each… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/robert-iang-ssekchy-dreva.dzintra_shultse_robertsik_inputrobertcantandodzintra_shultse_robertsik
