ThaiVanPhat95/wav2vec2-robust-uwb-atcosim-supcon-hybrid-4gram
Wav2Vec2 Robust UWB+ATCOSIM SupCon Hybrid with 4-Gram LM
Associated paper: Contrastive Regularization for Accent-Robust ASR.
This model performs English air-traffic-control automatic speech recognition. It combines a Wav2Vec2 CTC model trained with supervised contrastive learning and a 4-gram language model for decoding.
The hybrid training schedule combines:
- Similar-transcript SupCon batches
- Original and synthetic transcript pairs
- CTC-only batches
Intended Use
This model is intended for research on robust English ATC speech recognition. The included 4-gram language model is specialized for ATC terminology and phrase patterns.
Training Data
The acoustic model was trained using real UWB-ATCC speech, simulated ATCOSIM speech, and synthetic speech derived from their transcripts.
- UWB-ATCC: https://huggingface.co/datasets/Jzuluaga/uwb_atcc
- ATCOSIM: https://huggingface.co/datasets/Jzuluaga/atcosim_corpus
- Training code: https://github.com/thaivanphat95/robust-atc-asr
Users should review and comply with the original dataset terms before using or redistributing this model.
Usage
Install the language-model decoder dependency:
pip install pyctcdecodeimport torch
import soundfile as sf
from transformers import AutoModelForCTC, Wav2Vec2ProcessorWithLM
model_id = "thaivanphat95/wav2vec2-robust-uwb-atcosim-supcon-hybrid-4gram"
processor = Wav2Vec2ProcessorWithLM.from_pretrained(model_id)
model = AutoModelForCTC.from_pretrained(model_id).eval()
audio, sample_rate = sf.read("audio.wav")
inputs = processor(audio, sampling_rate=sample_rate, return_tensors="pt")
with torch.no_grad():
logits = model(input_values=inputs.input_values).logits
transcript = processor.decode(logits.cpu().numpy()[0]).text
print(transcript)Audio should be mono. Resample audio to 16 kHz before inference when necessary.
Citation
@article{thai2026contrastive,
title={Contrastive Regularization for Accent-Robust ASR},
author={Thai, Van-Phat and Dhruv, Aradhya and Pham, Duc-Thinh and Alam, Sameer},
journal={arXiv preprint arXiv:2605.03297},
year={2026},
doi={10.48550/arXiv.2605.03297}
}Limitations
- The acoustic and language models are specialized for English ATC speech.
- The language model may bias predictions toward common ATC phrase patterns.
- Performance may degrade for other domains, languages, or recording conditions.
- Model predictions should be reviewed before use in safety-critical settings.
License
The model weights and language model are not covered by the Apache License 2.0 used for the training code. Their use and redistribution may also be affected by the licenses and terms of the pretrained model, UWB-ATCC, ATCOSIM, source transcripts, and synthetic speech-generation systems.
