AdityaRaikar/mpse-smoe-speech-enhancement
042
MP-SENet + Supervised Mixture of Experts (S-MoE) for Dual-Bandwidth Speech Enhancement
Speech enhancement model handling both Narrowband (NB, 8kHz) and Wideband (WB, 16kHz) audio using Supervised Mixture of Experts.
Based on:
- [MP-SENet](https://arxiv.org/abs/2305.13686) — STFT-domain magnitude-phase speech enhancement backbone
- [S-MoE](https://arxiv.org/abs/2508.10009) — Supervised MoE with deterministic hard gating by bandwidth metadata
Pre-trained Baseline
The official pre-trained MP-SENet baseline (from yxlu-0102/MP-SENet) is included at pretrained/g_best_vb.pt:
- Trained on VoiceBank+DEMAND, 16kHz, 100 epochs
- 2.263M params, 247 state dict keys
- No need to train baseline from scratch — S-MoE training initializes directly from this checkpoint
Training (S-MoE only)
pip install -r requirements-gpu.txt
# S-MoE training (downloads pre-trained baseline automatically)
SMOE_EPOCHS=60 BATCH_SIZE=2 GRAD_ACCUM_STEPS=2 python train.pyThe script automatically:
- Downloads VoiceBank+DEMAND (~5GB)
- Pre-extracts WB/NB data into separate folders (no on-the-fly resampling)
- Downloads official pre-trained MP-SENet baseline from this repo
- Initializes S-MoE experts from baseline (shared → both WB+NB experts)
- Trains S-MoE on joint WB+NB data for 60 epochs
- Pushes trained model to HF Hub
Data Pipeline
All NB/WB data is pre-extracted to separate folders before training:
Architecture
Environment Variables
Files
References
@article{lu2023mpse,
title={MP-SENet: A Speech Enhancement Model with Parallel Denoising of Magnitude and Phase Spectra},
author={Lu, Ye-Xin and Ai, Yang and Ling, Zhen-Hua},
journal={arXiv preprint arXiv:2305.13686},
year={2023}
}