matbee/LFM2.5-Audio-1.5B-ONNX-cuda-gqa
06
README: full perf benchmarks (Q4 / hybrid / FP16) on RTX 4090 + Memcpy breakdown + VRAM-based config guide
Patch Q4 depthformer variants + rebuild Q4 unrolled_sample (clears cos/sin, sets do_rotary=0; explains residual Q4 MatMul CPU fallback)
Add reproduction scripts (patch_depthformer_gqa.py + unrolled builders) and rebuild instructions
Add unrolled+sample variants (in-graph Gumbel-max top-k), update NOTICE/README
Add files using upload-large-folder tool
initial commit
