ThakiCloud/Qwen3-Coder-30B-A3B-W4A16
1503
docs: 카드 메타데이터 보강 (base_model_relation/library_name/datasets/language)
Measure the Hopper case directly: W4A16 loses to FP8 there too (0.81-0.84x)
Answer the open kernel-tuning hedge with the four-way NVFP4 vs FP8 result; note the FP8 size discrepancy
card: NVFP4 sibling — 4-bit slowness is W4A16's, not 4-bit's
add model.safetensors
add tokenizer.json
add config.json
add qwen3coder_tool_parser.py
add chat_template.jinja
add tokenizer_config.json
add generation_config.json
add recipe.yaml
card: measured results before weights
initial commit
