CoolFace
Modelpublic

Luigi/moss-transcribe-diarize-zhtw-gguf

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
1likes1.9kdownloads
26 commits on main
3ee39822mo ago

README: q4mix-v2 as deployed weights (silence-robust QAT); document removed files

Luigi
bfeb7ab2mo ago

Remove deprecated files: zh-TW FT lineage (v5kl-v71), old moss-td-base-* conversions, superseded q4mix v1, unused campplus-cn-common

Luigi
805c42d2mo ago

q4mix-v2: silence-robust QAT (minimal-perturbation, 50 steps @1e-5 + L2 anchor on silence/patience/quiet pools). vs v1: silent-input marker-loop/garbage FIXED (deploy-env tail + held-out synth both 0 bytes); golden zh 89.25->96.65; 5-meeting WER/spk statistically tied (0.451/95.76 vs 0.450/95.58); golden en 98.73->97.22; -40dB sensitivity 96.3. Decoder tensors from QAT'd latents; encoder/adaptor/token_embd identical to v1.

Luigi
59391ef2mo ago

q4mix: uniform q4_k with token_embd held at f16 (0.76 GB, 51% smaller than q8mix). Validated under 90s windowing across 5 unseen AMI meetings: segmentation legitimacy identical (WRONG 24.8% vs q8mix 24.6%, +0.2pp), WER 0.1639 (matches q8mix exactly), speaker accuracy 99.0%, timestamp drift <=0.09s. Windowing is what makes q4 viable -- under one-pass decode q4 collapsed; per-window decodes are ~1000 tokens with a fresh KV cache, so quantization error has far less to accumulate over.

Luigi
0fa62752mo ago

README: reflect the purification-first pipeline; mark v5kl-v71 GGUFs as deprecated/unused

Luigi
f28ec5e2mo ago

mixed-precision q8_0: token_embd + qwen3 decoder held at f16 (plain q8_0 collapses zh long-form segmentation 312->69 utts; this fix restores 312/312, 100% text agreement vs f32, 1.55GB vs 3.64GB f32)

Luigi
5ca2fcf2mo ago

f32 GGUF converted from the official OpenMOSS checkpoint - the stage-1 parity artifact

Luigi
39d0b452mo ago

rehost campplus.gguf (was served from the WASM Space, which is being privatised)

Luigi
bf752792mo ago

base q5_k_m: stage-1 final. No regression vs q8/f16 on 32-min zh/en meetings, 0.80 vs 1.99 GB

Luigi
15bd08b2mo ago

base q8_0: accuracy-equivalent to f16 on 32-min zh/en meetings, 33% smaller

Luigi
4d2b6dc2mo ago

stage 1: base MOSS-TD f16 — numerically identical to f32 (17/17 segs, 1333 vs 1334 chars), half the size

Luigi
77709362mo ago

stage 1: base MOSS-TD q4_K_M — eviction-validated, 2x better general accuracy, 707MB

Luigi
a3a5a2b2mo ago

v7.1 f16 for inspection: MER 0.066 @3min vs 0.157 for q5 (q5 loses marker granularity)

Luigi
5cfcfed2mo ago

v7.1 q5_K_M: q4 collapses 3 speakers to 2; q5 preserves f16-identical diarization at 0.75GB

Luigi
51ccaaa2mo ago

README: v7 gguf is current (q4-robust QAT, repetition-loop fix); v6.1 marked superseded

Luigi
25a59f42mo ago

v7: CE-only QAT of v6.1 — restores q4 robustness (fixes the dense-speech repetition loop at the weight level; q4 ASCEND 0.352->0.317)

Luigi
7caada92mo ago

add CAM++ speaker embedding gguf (14MB) for cross-window linking — stable download for integrators

Luigi
3b238ce2mo ago

README: v6.1 gguf is current; full file table + updated comparison & benchmarks

Luigi
d6dc54e2mo ago

v6.1: marker-density FT (sentence-cadence time markers on dense speech; ASCEND all 0.285->0.267)

Luigi
1e890642mo ago

README: original-vs-ours comparison table

Luigi
4fedd3e2mo ago

v6-stream: streaming FT (45s bounded audio-KV) — CER 12.8 vs 27.2, linked DER parity, ASCEND all-bucket win

Luigi
43f10783mo ago

v5-kl QAT q4_k_m: QAT-defended weights survive q4 on long dense windows (plain-converted v5-kl q4 loses speaker tags / loops); pairs with RapidSpeech.cpp 193e0a9 decode fixes

Luigi
d346da33mo ago

Add embq8 variant (token_embd q8_0): enables all-GPU decode on Jetson Nano gen1 / sm_53 CUDA (no q6_K get_rows kernel) — 2.5x faster than the split; identical elsewhere

Luigi
f0c07473mo ago

Add model card

Luigi
c179b2f3mo ago

v5-kl MOSS-Transcribe-Diarize GGUF q4_k_m (RapidSpeech.cpp/ggml runtime)

Luigi
30ae1d53mo ago

initial commit

Luigi