cisco-ai/stupase-speech-enhancement
18
StuPASE: Studio-Quality Generative Speech Enhancement
This Space demonstrates StuPASE, a state-of-the-art generative speech enhancement model that removes noise and reverberation while preserving linguistic content and speaker identity, achieving studio-level perceptual quality.
How it works
Upload a noisy or reverberant speech recording (16 kHz mono recommended). StuPASE processes it through three stages:
- DeWavLM-R — Low-hallucination phonetic enhancement (fine-tuned from WavLM)
- CFM — Phonetic-guided acoustic enhancement via conditional flow matching
- Mel Vocoder — Reconstructs the enhanced waveform from mel features
Model
- Repository: cisco-ai/stupase
- Paper: StuPASE: Towards Low-Hallucination Studio-Quality Generative Speech Enhancement
- Code: cisco-open/pase
- License: Apache 2.0
