datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kinyarwanda_afrivoice_all_domains_v0.1
Afrivoice Kinyarwanda — All Domains
Combined dataset across 5 domains from the original source.
Note: the source dataset also includes a scripted_education domain, excluded
here due to a cluster of corrupted audio files in one of its shards.
Attribution
Original dataset: DigitalUmuganda/Afrivoice_Kinyarwanda
License: CC-BY-4.0
Attribution: Digital Umuganda
This dataset is derived from the above source and released under the same CC-BY-4.0 license.… See the full description on the dataset page: https://huggingface.co/datasets/ElizabethMwangi/kinyarwanda_afrivoice_all_domains_v0.1.kinyarwanda_afrivoice_all_domains_v0.2
Kinyarwanda AfriVoice — All Domains (v0.2)
Cleaned version of ElizabethMwangi/kinyarwanda_afrivoice_all_domains_v0.1.
Changes from v0.1
Removed rows with empty/null transcription values across all splits (train/validation/test)
Audio and domain labels unchanged; only null-transcription rows were dropped
Source
Original data from DigitalUmuganda/Afrivoice_Kinyarwanda (CC-BY-4.0),
extracted and concatenated across 5 domains (agriculture, education… See the full description on the dataset page: https://huggingface.co/datasets/ElizabethMwangi/kinyarwanda_afrivoice_all_domains_v0.2.ElizabethMidfordscotus-elizabeth_b_prelogar-audio
SCOTUS-sim audio: elizabeth_b_prelogar
Per-utterance audio clips from Oyez oral-argument mp3s, sliced at
the start_time / stop_time timestamps stored in the companion
scotus-sim/scotus-elizabeth_b_prelogar-training dataset.
Alignment
clip_NNNNN.wav in the tarball corresponds exactly to
audio_segments.jsonl[NNNNN] in the training companion dataset.
In metadata.jsonl each row carries the same 0-padded index in idx.
This supersedes the v1 tarball, which had systematic… See the full description on the dataset page: https://huggingface.co/datasets/scotus-sim/scotus-elizabeth_b_prelogar-audio.eliza-azheshka-kham-lika-ptashuk
Хам
Metadata
Author: Эліза Ажэшка
Title: Хам
Narrator: Ліка Пташук
Source Group: Аўдыёкнігі
Source: radiokultura.by
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders.
Target maximum split size: about 250 MB.… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/eliza-azheshka-kham-lika-ptashuk.eliza-azheshka-niziny-lika-ptashuk
Нізіны
Metadata
Author: Эліза Ажэшка
Title: Нізіны
Narrator: Ліка Пташук
Source Group: Аўдыёкнігі
Source: radiokultura.by
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders.
Target maximum split size: about… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/eliza-azheshka-niziny-lika-ptashuk.ElizabethOlsenknihi-be-eliza_azeska_nizinyknihi-be-eliza_azezka_chamEliza_Azeska_Niziny
