CoolFace
Modelpublic

Luigi/moss-transcribe-diarize-zhtw-gguf

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
1likes1.9kdownloads
Model Card

MOSS-Transcribe-Diarize GGUF weights

Status (2026-07-23): this repo serves a purification-first pipeline — the base OpenMOSS-Team/MOSS-Transcribe-Diarize (0.9B, Whisper-medium encoder + Qwen3-0.6B decoder, Apache-2.0), not fine-tuned, ported to C++/ggml and byte-verified against the genuine PyTorch reference before any optimization. Full writeup: `vieenrose/distil-vibevoice-asr`.

Files

filerole
moss-transcribe-base-q4mix-v2.ggufdeployed weights (759 MB) — mixed q4: every encoder/adaptor/decoder linear at Q4K, `tokenembd at f16, norms/biases f32. Decoder tensors come from a **silence-robust QAT** checkpoint (see below); encoder/adaptor/token_embd` are numerically identical to the plain base conversion
moss-transcribe-base-q8mix.ggufhigher-fidelity option (1.55 GB) — token_embd + full Qwen3 decoder at f16 (both measured Q80-collapse-sensitive), encoder + adaptor Q80. No silence QAT: its behavior on fully silent input has not been validated the way v2 has
moss-transcribe-base-f32.ggufthe byte-identity GATE reference (3.64 GB) — converted from the official checkpoint, f32, no fine-tuning
campplus.ggufCAM++ speaker embedding model, used for cross-window speaker linking in the demo

Why q4mix-v2 (silence robustness)

At f32 the model's EOS-on-silence decision has a thin margin (~+1.35 logits vs 15–17 on speech), so any q4 quantization noise could flip it: the previous q4mix free-ran marker loops ([0.00][S01][0.06]…) to the token budget on silent / unvoiced / near-silent audio. v2 fixes this in the weights — a minimal-perturbation QAT that trains only the silence decision class (silence negatives, garbage-recovery prefixes, quiet-lead-in patience composites, attenuated quiet-speech positives) with an L2 trust region to the base weights. Release validation vs the previous q4mix, identical harness:

  • —silent input → empty transcript in seconds (this is the correct no-speech result, not an error);
  • —zh 5-min golden agreement 89.3 → 96.7;
  • —5-meeting WER and speaker accuracy statistically unchanged;
  • —quiet-speech sensitivity preserved (−40 dB attenuated speech still transcribed);
  • —en single-pass golden −1.5 pt (97.2 vs 98.7) is the one known cost.

Engine and usage

Engine: `vieenrose/RapidSpeech.cpp` branch moss-pure, vendoring `localai-org/moss-transcribe.cpp` (MIT) unmodified. Live demo: `Luigi/moss-transcribe-diarize-cpp`.

Recommended flags: MT_KV_F16=1 MT_KV_EVICT_S=45 (near-lossless, large decode speedup on long audio). Integration details: integration note.

Removed files (2026-07-23)

The abandoned fine-tuned zh-TW lineage (moss-td-zhtw-v5kl…v71-*), the early base conversions (moss-td-base-*), the silence-fragile q4mix v1, and the unused campplus-cn-common.gguf were deleted from the tip of this repo. The fine-tuned line was measured to be over-specialised to its training domain and structurally fragile relative to the base model (structural-token logit margin 0.98 vs the base's 4.90 — quantization noise could flip marker decisions), which is why the project pivoted to the purified base. All removed files remain downloadable from this repo's git history (Files → History → pick a pre-2026-07-23 revision).