mlboydaisuke/qwen3.5-0.8B-CoreML
Card: base_model_relation: quantized (so the conversion lists under the base model's Quantizations, not Finetunes)
Link the card back to its collection and the request box
v1.8.0: full-vocab rep_penalty path — embed_weight.bin
v1.8.0: full-vocab rep_penalty path — weight.bin
v1.8.0: full-vocab rep_penalty path — model.mlmodel
v1.8.0: full-vocab rep_penalty path — Manifest.json
v1.8.0: full-vocab rep_penalty path — weight.bin
v1.8.0: full-vocab rep_penalty path — model.mlmodel
v1.8.0: full-vocab rep_penalty path — Manifest.json
v1.8.0: full-vocab rep_penalty path — weight.bin
v1.8.0: full-vocab rep_penalty path — model.mlmodel
v1.8.0: full-vocab rep_penalty path — Manifest.json
v1.8.0: full-vocab rep_penalty path — weight.bin
v1.8.0: full-vocab rep_penalty path — model.mlmodel
v1.8.0: full-vocab rep_penalty path — Manifest.json
Upload README.md with huggingface_hub
Upload README.md with huggingface_hub
docs: add file picker + standalone Python usage + tokenizer source
Add INT8 palettized decode (754 MB, 50% of fp16, parity preserved)
Initial upload: Qwen3.5-0.8B CoreML variants (decode + prefill + chunks + monolith)
initial commit
