haerik/seoul_corpus_asr_wavlm
haerik/seoulcorpusasr_wavlm
ESPnet2 ASR model for the Seoul Corpus, a Korean corpus of spontaneous interview speech (40 speakers, 1 hour each). Trained with the egs2/seoul_corpus/asr1 recipe.
The model transcribes speech and the corpus's non-speech annotations: <SIL>, <NOISE>, <VOCNOISE>, <LAUGH>, <UNKNOWN> and <PRIVATE.INFO> are targets, not discarded labels.
Architecture
WavLM-large (English-only pre-training) is used as a frozen SSL frontend, with a trainable E-Branchformer encoder and Transformer decoder on top (hybrid CTC/attention, ctc_weight 0.5).
Results
Test set: all 1409 annotated utterances. CER and WER are scored with the special tags removed (--nlsyms_txt), so they measure transcription rather than tag spelling; TER keeps them.
No language model: decoding is attention + CTC only (lm_weight 0.0).
Usage
from espnet2.bin.asr_inference import Speech2Text
speech2text = Speech2Text.from_pretrained("haerik/seoul_corpus_asr_wavlm")
text, *_ = speech2text(speech)[0]Corpus citation
Yun, Weonhee, Kyuchul Yoon, Sunwoo Park, Juhee Lee, Sungmoon Cho, Donghoon Kang, Koonhyuk Byun, Hyeonzu Hahn and Jungsun Kim. 2015. "The Korean Corpus of Spontaneous Speech." Phonetics and Speech Sciences 7(2), 103-109.
