logic65/Qwen3.6-Whittle-25B-A3B
card: next/ experimental qwen4exp build
next/: EXPERIMENTAL qwen4exp-format build (sigmoid GDN gate, identity HC, inert indexer, no n-gram yet). Runs on stock llama.cpp. GSM8K 86.5% via llama.cpp / 82.0% via transformers; parent 88.5%
Upload next/gate_act.json with huggingface_hub
Upload next/gsm8k_q4exp_stock_llamacpp.json with huggingface_hub
Upload next/gsm8k_sigmoid_transformers.json with huggingface_hub
next/: settled sigmoid-gate trainable state (routers, shared, LoRA, GDN gates) - experimental
card: support + authors
Upload eval/prune_report.json with huggingface_hub
Upload eval/gsm8k_base_qwen3.6-35b-a3b.json with huggingface_hub
Upload eval/gsm8k_pruned_healed.json with huggingface_hub
Qwen3.6-Whittle-25B-A3B: 256->180 experts, self-distilled heal (step 900), GSM8K 92.5%
initial commit
