flash-next
qwen3.8-flash-next-expert-traces
Qwen3.8-Flash-Next expert routing traces
Token-level routing traces of a deployed MoE model: for every token and every one of the
48 MoE layers, which experts the router chose, the top-32 router logits behind that choice,
and the exact hidden state the router read — plus, in v3, the state at many layers per token,
the post-final-norm state the LM head consumes, and the LM head's top-8 next-token candidates.
The corpus exists to answer one question: how well can the next tokens'… See the full description on the dataset page: https://huggingface.co/datasets/aswinkumar99/qwen3.8-flash-next-expert-traces.Qwen3.8-Flash-Next-GGUF-metricsqwen38-flash-next-int2-rtx5090-research
Qwen3.8-Flash-Next Mixed INT2 AutoRound on a Single RTX 5090
This benchmark and reproducibility artifact documents SGLang inference for the mixed-INT2 AutoRound Qwen3.8-Flash-Next checkpoint on one NVIDIA RTX 5090 Blackwell GPU. It covers a validated 256K / 262,144-token long context, MoE autotuning, tiered KV cache, CPU offload, and the device-local evidence showing why another Blackwell GPU's tuning configuration should not be copied blindly.
Headline inference… See the full description on the dataset page: https://huggingface.co/datasets/Jerrybro/qwen38-flash-next-int2-rtx5090-research.Qwen3.8-Flash-Next-expert-activation-map
Qwen3.8-Flash-Next Expert Activation Map
Per-(layer, expert) routing and output-importance statistics for
Qwen3.8-Flash-Next (Qwen4Exp architecture, 48 MoE layers x 512 routed
experts, top-k 10), measured on the unquantized bf16 checkpoint over a
2751-prompt, 28-domain calibration corpus.
The purpose is to answer, per layer, which experts carry the model's routed
output so that expert-level decisions (bf16 protection under quantization,
offload residency, pruning, warm-start… See the full description on the dataset page: https://huggingface.co/datasets/tcclaviger/Qwen3.8-Flash-Next-expert-activation-map.qwen3.8-flash-next-routing-traces
Qwen3.8 Flash Next Routing Traces
用于MoE路由、hidden预测及离线缓存分析的历史采集数据。不是模型权重,不是模型发布,也不是标准化性能基准。
本数据集配套Qwen3.8 Flash Next推理研究仓库。这里的模型名称沿用维护者的研究命名;不表示模型官方发布、官方认可或与其他同名模型版本兼容。数据来自基于llama.cpp深度修改的本地实验分支,不能假定上游llama.cpp可以直接重现全部采集行为。
归档已上传完成,Content-Length 与原文件一致,首尾片段校验相同;远端对象以SHA-256为存储键。
DATASET_INFO及归档内说明保留打包时的快照,其中"云盘链接尚未提供"等措辞已过期;当前分发地址以本页为准。
下载与核验
文件
用途
qwen3.8-hidden-routing-research.7z
5批原始采集与包内说明
SHA256SUMS.txt
压缩包SHA-256
archive-receipt.json… See the full description on the dataset page: https://huggingface.co/datasets/satsder/qwen3.8-flash-next-routing-traces.qwen38-flash-next-perfectblend-regen
