Luigi/x-asr-zh-en-streaming-zipformer2-gguf
0134
Nano benchmark: 4-engine cross-comparison + thread scaling (sherpa-onnx CUDA not runnable; RS CUDA+FP16 reaches sherpa-CPU parity)
Add measured on-device Jetson Nano (sm_53) benchmark: RS_GEMM_FP16 1.75x, token-exact
Update model card: add iq4_xs, per-weight accuracy table, document excluded sub-3-bit sweep
Add 960ms iq4_xs (lossless, 87MB)
Document q3_k (imatrix) near-lossless 69MB option
Add q3_k (imatrix, near-lossless, 69MB) + imatrix .dat calibration files
Document q8_0/q4_k as recommended formats + quant sweep guidance
Add Q8_0 variants (lossless: token-exact with f16, ~half size, ~1.2x faster on CPU)
Add all 4 streaming chunk variants (f16 + Q4_K) + model card
initial commit
