EmbeddedLLM/Inkling-Small-MXFP4
Inkling-Small-MXFP4
Model Overview
- Model architecture: Thinking Machines Lab Inkling-Small
- Parameters: 276B total / 12B active
- Input: Text, image, audio
- Output: Text
- Validated inference engine: vLLM
- Model optimizer: AMD Quark (
0.12.post1+rocm72.torch2.11) - Quantized layers: MoE routed experts in transformer layers 3 through 41
- Weight quantization: OCP MXFP4, static, group size 32, E8M0 scales
- Activation quantization: OCP MXFP4, dynamic, group size 32, E8M0 scales
This checkpoint was built from thinkingmachines/Inkling-Small revision b2d4f225a02032c5d154bff748ab5a00c5ca26e4 by applying AMD Quark OCP MXFP4 quantization to the BF16 routed experts. Routed-expert weights are stored as packed MXFP4 weights with E8M0 scales. Dense layers 0 through 2, attention, shared experts, embeddings, norms, the audio and vision towers, MTP, and other non-routed components remain in their source formats.
Environment
The file-to-file conversion and validation targeted AMD gfx950 and used:
- Container:
docker.io/rocm/vllm-dev:nightly_main_20260714 - Python: 3.12
- ROCm/HIP: 7.2
- PyTorch:
2.11.0+gitd0c8b1f - AMD Quark:
0.12.post1+rocm72.torch2.11 - Expert chunk size: 8
The paired quality evaluation used the same GPU type and TP8 topology, with Transformers 5.14.1 and vLLM commit 846e2d01a0be00acf31f1a354059c7c302c93042 (0.23.1rc1.dev1212+g846e2d01a).
Evaluation
Refer to the Inkling-Small model card for architecture, training, intended-use, safety, and acceptable-use details.
