Luigi/moss-transcribe-diarize-zhtw-gguf
README: q4mix-v2 as deployed weights (silence-robust QAT); document removed files
Remove deprecated files: zh-TW FT lineage (v5kl-v71), old moss-td-base-* conversions, superseded q4mix v1, unused campplus-cn-common
q4mix-v2: silence-robust QAT (minimal-perturbation, 50 steps @1e-5 + L2 anchor on silence/patience/quiet pools). vs v1: silent-input marker-loop/garbage FIXED (deploy-env tail + held-out synth both 0 bytes); golden zh 89.25->96.65; 5-meeting WER/spk statistically tied (0.451/95.76 vs 0.450/95.58); golden en 98.73->97.22; -40dB sensitivity 96.3. Decoder tensors from QAT'd latents; encoder/adaptor/token_embd identical to v1.
q4mix: uniform q4_k with token_embd held at f16 (0.76 GB, 51% smaller than q8mix). Validated under 90s windowing across 5 unseen AMI meetings: segmentation legitimacy identical (WRONG 24.8% vs q8mix 24.6%, +0.2pp), WER 0.1639 (matches q8mix exactly), speaker accuracy 99.0%, timestamp drift <=0.09s. Windowing is what makes q4 viable -- under one-pass decode q4 collapsed; per-window decodes are ~1000 tokens with a fresh KV cache, so quantization error has far less to accumulate over.
README: reflect the purification-first pipeline; mark v5kl-v71 GGUFs as deprecated/unused
mixed-precision q8_0: token_embd + qwen3 decoder held at f16 (plain q8_0 collapses zh long-form segmentation 312->69 utts; this fix restores 312/312, 100% text agreement vs f32, 1.55GB vs 3.64GB f32)
f32 GGUF converted from the official OpenMOSS checkpoint - the stage-1 parity artifact
rehost campplus.gguf (was served from the WASM Space, which is being privatised)
base q5_k_m: stage-1 final. No regression vs q8/f16 on 32-min zh/en meetings, 0.80 vs 1.99 GB
base q8_0: accuracy-equivalent to f16 on 32-min zh/en meetings, 33% smaller
stage 1: base MOSS-TD f16 — numerically identical to f32 (17/17 segs, 1333 vs 1334 chars), half the size
stage 1: base MOSS-TD q4_K_M — eviction-validated, 2x better general accuracy, 707MB
v7.1 f16 for inspection: MER 0.066 @3min vs 0.157 for q5 (q5 loses marker granularity)
v7.1 q5_K_M: q4 collapses 3 speakers to 2; q5 preserves f16-identical diarization at 0.75GB
README: v7 gguf is current (q4-robust QAT, repetition-loop fix); v6.1 marked superseded
v7: CE-only QAT of v6.1 — restores q4 robustness (fixes the dense-speech repetition loop at the weight level; q4 ASCEND 0.352->0.317)
add CAM++ speaker embedding gguf (14MB) for cross-window linking — stable download for integrators
README: v6.1 gguf is current; full file table + updated comparison & benchmarks
v6.1: marker-density FT (sentence-cadence time markers on dense speech; ASCEND all 0.285->0.267)
README: original-vs-ours comparison table
v6-stream: streaming FT (45s bounded audio-KV) — CER 12.8 vs 27.2, linked DER parity, ASCEND all-bucket win
v5-kl QAT q4_k_m: QAT-defended weights survive q4 on long dense windows (plain-converted v5-kl q4 loses speaker tags / loops); pairs with RapidSpeech.cpp 193e0a9 decode fixes
Add embq8 variant (token_embd q8_0): enables all-GPU decode on Jetson Nano gen1 / sm_53 CUDA (no q6_K get_rows kernel) — 2.5x faster than the split; identical elsewhere
Add model card
v5-kl MOSS-Transcribe-Diarize GGUF q4_k_m (RapidSpeech.cpp/ggml runtime)
initial commit
