twangodev/radiotalk-us-audio-grok-noisy
radiotalk-us-audio-grok-noisy VHF-AM channel-degraded counterpart to twangodev/radiotalk-us-audio-grok-clean: 3 independently-degraded variants per clean utterance (bandpass, noise, fading, heterodyne, PTT clicks, codec artifacts — the same radiotalk radio pipeline behind the higgs/tada noisy sets). 1,583,103 rows covering all rendered scenarios of twangodev/radiotalk-us-transcripts-grok-4.20-50k. Difficulty (Grok STT) On a 10,000-utterance sample, Grok STT scores… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-audio-grok-noisy.
radiotalk-us-audio-grok-noisy
VHF-AM channel-degraded counterpart to `twangodev/radiotalk-us-audio-grok-clean`: 3 independently-degraded variants per clean utterance (bandpass, noise, fading, heterodyne, PTT clicks, codec artifacts — the same radiotalk radio pipeline behind the higgs/tada noisy sets). 1,583,103 rows covering all rendered scenarios of `twangodev/radiotalk-us-transcripts-grok-4.20-50k`.
Difficulty (Grok STT)
On a 10,000-utterance sample, Grok STT scores 41.8% mean WER / 22% exact-match here vs 5.8% / 75% on the clean set — the VHF-AM channel simulation is doing its job as a robustness stressor.
CC-BY-4.0. Attribute the radiotalk project + xAI for the underlying TTS.
