datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
live-translator-packs
Live Translator — Offline Model Packs
On-demand offline model packs for the Live Translator app. Downloaded per language pair;
the shared common/ bundle installs once on the first pair.
Layout
common/ # downloaded ONCE, on the first pair
asr/ # Whisper tiny int8 (multilingual ASR + language id), sherpa-onnx
whisper-encoder.onnx
whisper-decoder.onnx
whisper-tokens.txt
vad/silero_vad.onnx # Silero… See the full description on the dataset page: https://huggingface.co/datasets/Gstam21/live-translator-packs.kumu-livestream-segmented
halo-livestream
Real Taglish code-switching from livestreams — every segment carries forced-alignment confidence, ASR round-trip CER, SNR, loudness and overlap flags.
62 segments · 3 speakers · seed release
🌱 This is a seed release — 62 segments, about 7 minutes
It exists to publish the pipeline and the schema, not to be a training
corpus. Nothing here is big enough to train on. What is worth your time is the
per-segment quality metadata below — and the… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/kumu-livestream-segmented.
