CoolFace
Modelpublic

neurlang/ipa-whipstr-base-48khz-cv-21

sourceHugging Facegpl-2.0updated 3mo agoView on Hugging Face
2likes17downloads
Model Card

Neurlang Whipstr STT (ASR)

A deep learning automatic speech recognition (ASR) system for transcribing speech audio into IPA text using transformer-based sequence-to-sequence models.

  • —Language: Universal (IPA), 74+ languages
  • —Model Github: neurlang/whipstr https://github.com/neurlang/whipstr
  • —Model Dataset: Common Voice 21
  • —Model-Native Sample Rates: 8000 Hz, 16000 Hz, 24000 Hz, 32000 Hz, 48000 Hz
  • —Degraded-Performance Sample Rates: 11025 Hz, 22050 Hz, 44100 Hz
  • —License: GPL v2
  • —Release: 2026-07-07
  • —Size: 186 MB
  • —Total parameters:
  • —Encoder: 7 220 576
  • —Transformer: 7 537 184
  • —Total: 14 757 760
  • —CER: 48.68% (51.32% success rate)
  • —Note: Averaged across all supported languages, works better on higher resource languages
  • —WER: 92.58% (7.42% success rate)
  • —Note: Averaged across all supported languages, works better on higher resource languages
  • —Training Details:
  • —Hardware: Nvidia Spark
  • —Batch size: 1
  • —Samples: 960000
  • —Duration: 1:11:10:00
  • —Runs: Jul 5 14:58 - Jul 5 17:45, Jul 5 22:41 - Jul 6 04:38, Jul 6 04:46 - Jul 7 07:12 (shut down at Jul 7 07:36)

Inference code

bash
git clone https://github.com/neurlang/whipstr.git
cd whipstr/
uv run --with torch --with transformers --with phase-spectrogram stt_infer_hf.py --audio /home/m/Downloads/LJ001-0001.wav --model neurlang/ipa-whipstr-base-48khz-cv-21

Output:

bash
Loading weights: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 139/139 [00:00<00:00, 13371.75it/s]
Transcription: ˈɪt ˈæt juː pˈæst ðə kˈɔldz tˈɑɹk fˈeɪs vˈæst ðə ɹˈɛpɪtʃən lˈupənəl ˈi ˈæz bɪhˈeɪviɚ wˈi sˈɑloʊks lˈɛkəmˌɑɹəl hˈæn jˈæt lˈɚnd bˈeɪsɪks sˈɛktɪs ˈækmɪkmɪkəɹɪkənɪkəksɪksɪksɪksɪksɪksɪksɪksɪkstəkstəksɪksɪk

Explaination:

IPAGround truthComment
ˈɪtItExact match.
ˈæt juːgets you/ɡɛts/ was apparently lost, leaving something like "at you."
pˈæstpastGood match.
ðətheExact.
kˈɔldzcold-Reasonable; final /z/ is likely a transcription artifact.
tˈɑɹkstartThe /s/ was dropped and /st/ became /k/, a common recognition error in noisy speech.
fˈeɪsphaseVery close.
vˈæstfast/f/→/v/ substitution.
ðəTheExact.
ɹˈɛpɪtʃənrepetitionClearly recognizable.
lˈupənəlloop / no-EOSThis portion collapsed badly; "loop no-EOS" became something like "lupənəl."
ˈi ˈæzbehavior (beginning omitted)The start of "behavior" appears to have shifted.
bɪhˈeɪviɚbehaviorExcellent match.
wˈiweExact.
sˈɑloʊkssaw looksThis is actually close to "saw looks," with the pause removed.
lˈɛkəmˌɑɹəllike the modelQuite distorted, but you can hear the rough rhythm.
hˈænhadn'tThe final /t/ and /d/ were lost.
jˈætyetVery close.
lˈɚndlearnedExact.
bˈeɪsɪksbasicEssentially correct.
sˈɛktɪsseq2seqThis was the biggest failure—only the initial /sɛk/ of "seq2seq" survived, while
loopsmechanics"mechanics" disappeared entirely.

End of training data:

Step 960000 is chosen as this release

Step**WER****CER**
94300094.21%49.58%
94400093.16%49.96%
94500094.09%50.22%
94600095.25%50.32%
94700096.18%50.30%
94800093.97%50.78%
94900092.93%49.63%
95000094.55%50.84%
95100094.90%50.11%
95200092.35%49.85%
95300093.74%49.08%
95400093.74%49.80%
95500092.70%49.87%
95600094.32%50.37%
95700094.21%48.90%
95800093.40%50.08%
95900092.82%50.20%
96000092.58%48.68%
96100095.13%50.32%
96200092.93%49.56%
96300094.21%50.84%
96400095.48%50.44%
96500094.32%50.39%
96600093.05%50.72%
96700093.97%51.04%
96800091.19%49.94%
96900095.60%50.82%
97000096.52%50.16%
97100092.82%51.22%