Remlemon/yourmt3-vocal-enhanced
1211
A MIDI transcription model focuses on vocal, fine-tuned on the MIR-ST500 Chinese vocal.
Files
model.pt- inference weights.config.json- model/audio/task configuration.r47b_infer.py- local vocal stem to MIDI CLI.runtime_deps/,yourmt3_train_src/- minimal runtime source.
Usage
pip install -r requirements.txt
python r47b_infer.py vocals.wav --output-midi vocals.mid --device cudaInput should be a separated vocal stem.
Metrics
Metrics are reported on MIR-ST500 vocal transcription splits; this model repo does not include dataset audio or labels.
