P2Enjoy/VibeVoice-ASR-BitNet-slim
Long-form confirmed on a second TED talk: parity with whisper-small
Final: weights-focused card, real long-form results, documented limits
Correction: long-form is this model's strength, not its weakness (WER 8.88 vs whisper-small 27.57)
Long-form test results: retract unverified single-pass and speaker claims
Final benchmark table: WER + speed over upstream baseline, all engines
Residual-fusion default: +0.04 WER row, engine ratio ~3.0x upstream
Accuracy in the open: four-engine WER table incl. whisper; speed as engine ratios
Refresh accuracy tables on the current portable build (094b53c): 4/6 languages identical, corpus +0.30; add MLS-FR anchor row and hotword biasing note
Final results: same-build WER comparison (corpus 13.65 vs 13.95), deterministic runtime, speed pointer to repo
Split responsibilities with the GitHub README: card owns model+WER, repo owns speed+reproduction
Add measured WER cost of dropping the F16 output projection
Upload vibeasr-vae-encoder-i8_s.gguf with huggingface_hub
Upload vibeasr-lm-i2_s-tied.gguf with huggingface_hub
Upload vocab.json with huggingface_hub
Upload tokenizer_config.json with huggingface_hub
Upload tokenizer.json with huggingface_hub
Upload config.json with huggingface_hub
Upload README.md with huggingface_hub
initial commit
