dominant-strategies/quai-deepseek-v4-flash-igemm-w8a8
Quai Inference DeepSeek-V4-Flash IGEMM W8A8
Built for Quai Inference.
This repository contains InferenceGemm/Tensor Work Proof verifier-facing checkpoint artifacts derived from deepseek-ai/DeepSeek-V4-Flash. It is not an official DeepSeek release.
Artifact precision: W8A8 int8 weights derived from DeepSeek-V4-Flash FP4 experts, FP8 non-expert weights, and BF16 router/native tensors. Artifact format: igemm-w8a8-v2.
Contents
igemm-*.safetensors: Quai receipt-native tensor payloads.manifest.json: verifier-facing tensor manifest and Merkle roots.index.json: checkpoint index for the quantized tensor shards.runtime-alias-audit.json, when present: runtime-to-manifest alias audit.- DeepSeek FP4/FP8 manifest audit and benchmark evidence, when present.
benchmark-evidence/preflight/, when present: preflight serving and receipt-path smoke evidence.docs/, when present: Quai production receipt plan/status notes.- tokenizer/config files copied from the base model directory.
Benchmark Status
Preproduction DeepSeek-V4-Flash W8A8 checkpoint artifact. This is the consensus-first W8A8 path derived from the upstream FP4/FP8 checkpoint. Production SGLang W8A8 receipt benchmark evidence, strict receipt smoke, and Go verifier promotion are still pending.
The paper benchmark rows were produced with strict SGLang Tensor Work Receipt emission and Go verification. The raw benchmark logs, receipts, and local evidence packet are intentionally not uploaded to this model repository.
Loading
These files are intended for the Quai InferenceGemm harness in this repository, not vanilla transformers weight loading:
checkpoints/<this-checkpoint>/Use the base model tokenizer/config with the InferenceGemm quantized payloads and verifier manifest.
Upstream
- Base model:
deepseek-ai/DeepSeek-V4-Flash - Upstream license:
mit
Redistribution must comply with the upstream model license and applicable export control restrictions.
