katostrofik/qwen35moe-blackhole-p100-slim-gate-up-fused-2026-05-24
0
Qwen35MoE Blackhole P100 Slim Gate-Up Fused Artifact
This artifact is the large-file companion to the GitHub handoff repo/branch for the single-card Qwen 3.5/Qwen3.5-MoE-family Blackhole P100 bring-up.
It contains the TTNN tensorbin payload needed by:
python llm/qwen35moe_resident_cache_load_walk.py \
--bundle llm/cache/qwen36_resident_cache/full_resident_p100_slim_gate_up_fused_bundle.json \
--streaming
python llm/qwen35moe_showable_baseline.py \
--bundle llm/cache/qwen36_resident_cache/full_resident_p100_slim_gate_up_fused_bundle.json \
--tokenizer-json llm/models/Qwen3.6-35B-A3B-FP8/tokenizer.jsonImportant Claim
This was validated locally on one Blackhole P150A. The slim resident payload and runtime estimate should fit a 28 GiB P100, but direct P100 validation is still pending.
Contents
Required model/cache payload:
llm/cache/qwen36_resident_cache/full_resident_p100_slim_gate_up_fused_bundle.jsonllm/cache/qwen36_resident_cache/full_resident_p100_slim_gate_up_fused_report.jsonllm/cache/qwen36_resident_cache/materialized_mixed_fit_dense_allllm/cache/qwen36_resident_cache/materialized_mixed_fit_globalllm/cache/qwen36_resident_cache/materialized_mixed_fit_layer0_routed_banksllm/cache/qwen36_resident_cache/materialized_mixed_fit_routed_banks_layers1_39llm/cache/qwen36_resident_cache/materialized_mixed_fit_routed_gate_up_fusedllm/cache/qwen36_resident_cache/materialized_mixed_fit_routers_allllm/cache/qwen36_resident_cache/materialized_mixed_fit_shared_experts_allllm/cache/qwen36_resident_cache/global_weights_bf16_embeddingllm/models/Qwen3.6-35B-A3B-FP8/tokenizer.json
Reference outputs:
llm/cache/qwen36_resident_cache/full_resident_p100_slim_gate_up_fused_load_walk_streaming.jsonllm/cache/qwen36_resident_cache/full_resident_p100_slim_gate_up_fused_load_walk_resident.jsonllm/cache/qwen36_resident_cache/qwen35moe_p100_slim_showable_baseline.json
Size
- Slim resident bundle tensorbins: about
18.81 GiB - BF16 embedding tensor: about
0.95 GiB - Tokenizer JSON: about
12 MB - Estimated loaded runtime, including embedding and decode state: about
19.85 GiB
Expected Baseline Output
Local AI empowers users to run machine learning models directly on their own devices, ensuring data privacy and eliminating reliance on external cloud servers.<|im_end|>