xpuenabler/gpt-oss-15.5b-23E-SFT-v6-gpu
03
v6-gpu: v5-gpu recipe + Expert ReduceSum decomposition baked into the export. INT4 g=64 attn/dense/lm_head/embed; MoE experts and router FP16; SDPA + Softmax + Expert-ReduceSum decomposed; dtype=fp16 + defensive fp32->fp16 rewrite. Stateful, GenAI LLMPipeline-compatible.
initial commit
