CoolFace
Modelpublic

katostrofik/qwen35moe-blackhole-p100-slim-gate-up-fused-2026-05-24

sourceHugging Faceupdated 4mo agoView on Hugging Face
0likes
Model Card

Qwen35MoE Blackhole P100 Slim Gate-Up Fused Artifact

This artifact is the large-file companion to the GitHub handoff repo/branch for the single-card Qwen 3.5/Qwen3.5-MoE-family Blackhole P100 bring-up.

It contains the TTNN tensorbin payload needed by:

bash
python llm/qwen35moe_resident_cache_load_walk.py \
  --bundle llm/cache/qwen36_resident_cache/full_resident_p100_slim_gate_up_fused_bundle.json \
  --streaming

python llm/qwen35moe_showable_baseline.py \
  --bundle llm/cache/qwen36_resident_cache/full_resident_p100_slim_gate_up_fused_bundle.json \
  --tokenizer-json llm/models/Qwen3.6-35B-A3B-FP8/tokenizer.json

Important Claim

This was validated locally on one Blackhole P150A. The slim resident payload and runtime estimate should fit a 28 GiB P100, but direct P100 validation is still pending.

Contents

Required model/cache payload:

  • —llm/cache/qwen36_resident_cache/full_resident_p100_slim_gate_up_fused_bundle.json
  • —llm/cache/qwen36_resident_cache/full_resident_p100_slim_gate_up_fused_report.json
  • —llm/cache/qwen36_resident_cache/materialized_mixed_fit_dense_all
  • —llm/cache/qwen36_resident_cache/materialized_mixed_fit_global
  • —llm/cache/qwen36_resident_cache/materialized_mixed_fit_layer0_routed_banks
  • —llm/cache/qwen36_resident_cache/materialized_mixed_fit_routed_banks_layers1_39
  • —llm/cache/qwen36_resident_cache/materialized_mixed_fit_routed_gate_up_fused
  • —llm/cache/qwen36_resident_cache/materialized_mixed_fit_routers_all
  • —llm/cache/qwen36_resident_cache/materialized_mixed_fit_shared_experts_all
  • —llm/cache/qwen36_resident_cache/global_weights_bf16_embedding
  • —llm/models/Qwen3.6-35B-A3B-FP8/tokenizer.json

Reference outputs:

  • —llm/cache/qwen36_resident_cache/full_resident_p100_slim_gate_up_fused_load_walk_streaming.json
  • —llm/cache/qwen36_resident_cache/full_resident_p100_slim_gate_up_fused_load_walk_resident.json
  • —llm/cache/qwen36_resident_cache/qwen35moe_p100_slim_showable_baseline.json

Size

  • —Slim resident bundle tensorbins: about 18.81 GiB
  • —BF16 embedding tensor: about 0.95 GiB
  • —Tokenizer JSON: about 12 MB
  • —Estimated loaded runtime, including embedding and decode state: about 19.85 GiB

Expected Baseline Output

text
Local AI empowers users to run machine learning models directly on their own devices, ensuring data privacy and eliminating reliance on external cloud servers.<|im_end|>