CoolFace
Modelpublic

ThaiVanPhat95/wav2vec2-robust-uwb-atcosim-supcon-hybrid-4gram

sourceHugging Faceotherupdated 4mo agoView on Hugging Face
0likes11downloads
Model Card

Wav2Vec2 Robust UWB+ATCOSIM SupCon Hybrid with 4-Gram LM

Associated paper: Contrastive Regularization for Accent-Robust ASR.

This model performs English air-traffic-control automatic speech recognition. It combines a Wav2Vec2 CTC model trained with supervised contrastive learning and a 4-gram language model for decoding.

The hybrid training schedule combines:

  • —Similar-transcript SupCon batches
  • —Original and synthetic transcript pairs
  • —CTC-only batches

Intended Use

This model is intended for research on robust English ATC speech recognition. The included 4-gram language model is specialized for ATC terminology and phrase patterns.

Training Data

The acoustic model was trained using real UWB-ATCC speech, simulated ATCOSIM speech, and synthetic speech derived from their transcripts.

  • —UWB-ATCC: https://huggingface.co/datasets/Jzuluaga/uwb_atcc
  • —ATCOSIM: https://huggingface.co/datasets/Jzuluaga/atcosim_corpus
  • —Training code: https://github.com/thaivanphat95/robust-atc-asr

Users should review and comply with the original dataset terms before using or redistributing this model.

Usage

Install the language-model decoder dependency:

bash
pip install pyctcdecode
python
import torch
import soundfile as sf
from transformers import AutoModelForCTC, Wav2Vec2ProcessorWithLM

model_id = "thaivanphat95/wav2vec2-robust-uwb-atcosim-supcon-hybrid-4gram"

processor = Wav2Vec2ProcessorWithLM.from_pretrained(model_id)
model = AutoModelForCTC.from_pretrained(model_id).eval()

audio, sample_rate = sf.read("audio.wav")
inputs = processor(audio, sampling_rate=sample_rate, return_tensors="pt")

with torch.no_grad():
    logits = model(input_values=inputs.input_values).logits

transcript = processor.decode(logits.cpu().numpy()[0]).text
print(transcript)

Audio should be mono. Resample audio to 16 kHz before inference when necessary.

Citation

bibtex
@article{thai2026contrastive,
  title={Contrastive Regularization for Accent-Robust ASR},
  author={Thai, Van-Phat and Dhruv, Aradhya and Pham, Duc-Thinh and Alam, Sameer},
  journal={arXiv preprint arXiv:2605.03297},
  year={2026},
  doi={10.48550/arXiv.2605.03297}
}

Limitations

  • —The acoustic and language models are specialized for English ATC speech.
  • —The language model may bias predictions toward common ATC phrase patterns.
  • —Performance may degrade for other domains, languages, or recording conditions.
  • —Model predictions should be reviewed before use in safety-critical settings.

License

The model weights and language model are not covered by the Apache License 2.0 used for the training code. Their use and redistribution may also be affected by the licenses and terms of the pretrained model, UWB-ATCC, ATCOSIM, source transcripts, and synthetic speech-generation systems.