CoolFace
Modelpublic

camillebrl/whisper-large-v3-synthetic-20k-matched-psd_snrm10_drv6_o4

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes98downloads
Model Card

whisper-large-v3 fine-tuned on fully synthetic ATC speech (STEAK v2)

openai/whisper-large-v3 fine-tuned only on synthetic air-traffic-control speech — no real ATC audio was used for training. It nonetheless beats a real-ATCO2 fine-tune on two of three real European test sets.

Training data (all parameters)

  • —Corpus: STEAK v2, fully synthetic. Texts generated from a 119-command-type ATC-exchange ontology (scraped European registries: 8,147 airline telephonies, 11,134 waypoints, ~2,500 controller stations), voiced by LongCat-AudioDiT-3.5B voice cloning on verified, TF-IDF-matched LiveATC prompt clips, then passed through a parametric VHF-noise chain `psd_snrm10_drv6_o4` (PSD noise @ SNR −10 dB → soft-clip drive 6 → Butterworth order 4, 250–3600 Hz, 16 kHz).
  • —Prompt pairing: matched (each generated text is paired with the closest verified LiveATC clip by cosine TF-IDF, Hungarian assignment without replacement).
  • —Volume / budget: 20,000 distinct utterances seen once (1 pass), drawn from a 110,000-utterance matched pool — 1,250 optimiser steps, effective batch 16. This early-stop point is the best checkpoint: a full pass (6,876 steps) overfits the synthesis distribution and degrades on the real test sets.

Evaluation — WER (%) on real European test sets

test setglobal WERcallsigncommandvalue
airbus (French, TLS/BOD approach, 3,589 utt.)31.936.212.914.7
atco2 (Czech/multinational, 200 held-out)30.631.128.219.4
uwb (Czech, en-route, 800)41.133.123.423.2

Extraction recall on airbus (%): command 88.5 · value 86.1 · telephony 64.4 · callsign 51.6.

For reference, a whisper-large-v3 fine-tuned on real ATCO2 audio at the same optimiser budget scores 36.4 / 30.2 / 39.6 global WER on airbus / atco2 / uwb — so this synthetic-only model is 4.5 WER points better on airbus and on par on atco2, with no real ATC training audio.

Usage

python
from transformers import pipeline
asr = pipeline("automatic-speech-recognition",
               model="camillebrl/whisper-large-v3-synthetic-20k-matched-psd_snrm10_drv6_o4")
print(asr("atc_clip.wav", generate_kwargs={"language": "en", "task": "transcribe"}))