datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
liveatc
Dataset Card for LiveATC Recordings (Partial 2024-08-26)
Dataset Summary
This dataset contains a partial collection of 21,172 air traffic control audio recordings from LiveATC.net for the date August 26, 2024. The recordings are organized by ICAO airport code and stored in .tar.zst archives. Due to rate limiting on LiveATC.net, this dataset represents an incomplete sample rather than a full day's recordings as originally planned. It is provided as-is for research… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/liveatc.live-translator-packs
Live Translator — Offline Model Packs
On-demand offline model packs for the Live Translator app. Downloaded per language pair;
the shared common/ bundle installs once on the first pair.
Layout
common/ # downloaded ONCE, on the first pair
asr/ # Whisper tiny int8 (multilingual ASR + language id), sherpa-onnx
whisper-encoder.onnx
whisper-decoder.onnx
whisper-tokens.txt
vad/silero_vad.onnx # Silero… See the full description on the dataset page: https://huggingface.co/datasets/Gstam21/live-translator-packs.kumu-livestream-raw
halo-livestream-raw
Unsegmented source recordings behind sapinsapin/halo-livestream — the archival input to the pipeline, not a training set.
1 recording(s) · 26:56 · 9.0MB
🔒 Gated on purpose
Full-length conversation between named, identifiable speakers is a very
different privacy proposition from the short segments in the processed
dataset, so access here is gated: request it and agree to the terms above.
If what you want is segmented, quality-scored… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/kumu-livestream-raw.kumu-livestream-segmented
halo-livestream
Real Taglish code-switching from livestreams — every segment carries forced-alignment confidence, ASR round-trip CER, SNR, loudness and overlap flags.
62 segments · 3 speakers · seed release
🌱 This is a seed release — 62 segments, about 7 minutes
It exists to publish the pipeline and the schema, not to be a training
corpus. Nothing here is big enough to train on. What is worth your time is the
per-segment quality metadata below — and the… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/kumu-livestream-segmented.
