datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kumu-livestream-raw
halo-livestream-raw
Unsegmented source recordings behind sapinsapin/halo-livestream — the archival input to the pipeline, not a training set.
1 recording(s) · 26:56 · 9.0MB
🔒 Gated on purpose
Full-length conversation between named, identifiable speakers is a very
different privacy proposition from the short segments in the processed
dataset, so access here is gated: request it and agree to the terms above.
If what you want is segmented, quality-scored… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/kumu-livestream-raw.kumu-livestream-segmented
halo-livestream
Real Taglish code-switching from livestreams — every segment carries forced-alignment confidence, ASR round-trip CER, SNR, loudness and overlap flags.
62 segments · 3 speakers · seed release
🌱 This is a seed release — 62 segments, about 7 minutes
It exists to publish the pipeline and the schema, not to be a training
corpus. Nothing here is big enough to train on. What is worth your time is the
per-segment quality metadata below — and the… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/kumu-livestream-segmented.
