CoolFace
Modelpublic

haerik/seoul_corpus_asr_wavlm

sourceHugging Faceapache-2.0updated 4d agoView on Hugging Face
0likes
Model Card

haerik/seoulcorpusasr_wavlm

ESPnet2 ASR model for the Seoul Corpus, a Korean corpus of spontaneous interview speech (40 speakers, 1 hour each). Trained with the egs2/seoul_corpus/asr1 recipe.

The model transcribes speech and the corpus's non-speech annotations: <SIL>, <NOISE>, <VOCNOISE>, <LAUGH>, <UNKNOWN> and <PRIVATE.INFO> are targets, not discarded labels.

Architecture

WavLM-large (English-only pre-training) is used as a frozen SSL frontend, with a trainable E-Branchformer encoder and Transformer decoder on top (hybrid CTC/attention, ctc_weight 0.5).

total parameters362.76 M
trainable47.30 M (13.0%)
training data7238 utterances, 26.30 h
epochs40

Results

Test set: all 1409 annotated utterances. CER and WER are scored with the special tags removed (--nlsyms_txt), so they measure transcription rather than tag spelling; TER keeps them.

CERWERTER
13.436.723.4

No language model: decoding is attention + CTC only (lm_weight 0.0).

Usage

python
from espnet2.bin.asr_inference import Speech2Text

speech2text = Speech2Text.from_pretrained("haerik/seoul_corpus_asr_wavlm")
text, *_ = speech2text(speech)[0]

Corpus citation

Yun, Weonhee, Kyuchul Yoon, Sunwoo Park, Juhee Lee, Sungmoon Cho, Donghoon Kang, Koonhyuk Byun, Hyeonzu Hahn and Jungsun Kim. 2015. "The Korean Corpus of Spontaneous Speech." Phonetics and Speech Sciences 7(2), 103-109.