Koshkasa/Vortex5_Shadow-Siren-26B-A4B-MXFP4_MOE-GGUF
What's that?
MXFP4_MOE quantization of Vortex5/Shadow-Siren-26B-A4B
Use case
Running the model on consumer-grade Blackwell GPUs (due to native fp4 support). Running on other hardware IS POSSIBLE. I lack conclusive data to say if it's a good idea.
Benchmarks
DISCLAIMER: benchmarking conducted on another G4-26B-A4B merge. I expect results to be interchangeable given the architecture.
tested on: --temp 0 -ngl 999 -b 512 -ub 512 -ctk q80 -ctv q80 pp65536/tg1024 test for Q4KS was run at -ngl 29 instead - it did not fit otherwise. pp65536/tg1024 test for MXFP4_MOE(16 exps) was run at -ngl 30 instead - it did not fit otherwise.
Conclusion
MXFP4MOE allows either Q4KS intelligence at -16.5% wall time, or higher than Q4K_S intelligence at the cost of +5.43% wall time (with experts/tok override to 16), while being of smaller file size. Is it worth it? I don't know. At this point, my particular config isn't compute-limited, but rather bottlenecked by PCI-E. WYSIWYG.
Disclosure
My only contribution is compute. This is not my merge. Have fun.
Cheers
[Google](https://huggingface.co/google) - the base model. [Vortex5](https://huggingface.co/Vortex5) - for the merge effort. Everyone whose finetunes were included in the merge!
