int2
gemma-4-26B-A4B-it-uncensored-abliterix-MLX-2bit-int2-affineQwen3.8-Flash-Next-int2-mixed-AutoRound-24GB-SGLangAlphaAI-Chatty-INT2-i1-GGUFAlphaAI-Chatty-INT2-GGUFQwen3.8-27B-INT8-W8A16-MTP-Q4_K_M-GGUFQwen3.5-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking-MLX-2bit-int2-affineDeepSeek-R1-Distill-Qwen-32B-DASHQ-INT2-g32phi-4-DASHQ-INT2-g32
Datasets
All datasets matching “int2”qwen38-flash-next-int2-rtx5090-research
Qwen3.8-Flash-Next Mixed INT2 AutoRound on a Single RTX 5090
This benchmark and reproducibility artifact documents SGLang inference for the mixed-INT2 AutoRound Qwen3.8-Flash-Next checkpoint on one NVIDIA RTX 5090 Blackwell GPU. It covers a validated 256K / 262,144-token long context, MoE autotuning, tiered KV cache, CPU offload, and the device-local evidence showing why another Blackwell GPU's tuning configuration should not be copied blindly.
Headline inference… See the full description on the dataset page: https://huggingface.co/datasets/Jerrybro/qwen38-flash-next-int2-rtx5090-research.INT2
Just 3 numbers ¯_(ツ)_/¯
enbeds_for_int20hint2
