CoolFace
Modelpublic

0xSero/DeepSeek-V4-Flash-180B-GGUF

sourceHugging Facemitupdated 4mo agoView on Hugging Face
14likes140downloads
Model Card
[!TIP] [Support this work →](https://donate.sybilsolutions.ai) · X · GitHub · REAP paper · Cerebras REAP

DeepSeek-V4-Flash-180B-GGUF

GGUF quantization of 0xSero/DeepSeek-V4-Flash-180B.

At a glance

Base model0xSero/DeepSeek-V4-Flash-180B
FormatGGUF
Total params180B
Active / token—
Experts / layer—
Layers—
Hidden size—
Context—
On-disk size164 GB

Which variant should I pick?

VariantFormatLink
DeepSeek-V4-Flash-162BBF16link
DeepSeek-V4-Flash-162B-GGUFGGUFlink
DeepSeek-V4-Flash-180BBF16link
DeepSeek-V4-Flash-180B-GGUF (this)GGUFlink
DeepSeek-V4-Flash-213BBF16link

This repository contains DS4/DwarfStar GGUF conversions of DeepSeek-V4-Flash-Spark.

The GGUFs point back to the original Spark Hugging Face model:

  • —Original Spark model: https://huggingface.co/0xSero/DeepSeek-V4-Flash-180B
  • —Conversion source checkpoint: https://huggingface.co/0xSero/DeepSeek-V4-Flash-180B-codex-K160-REAP
  • —Runtime/converter repo: https://github.com/antirez/ds4
  • —Spark deployment repo: https://github.com/0xSero/deepseek-spark

Files

FileSizeSHA256
DeepSeek-V4-Flash-Spark-Q2-REAP-ds4.gguf53.52 GiBdae2ed196e8ad87d6667d3fa04f65d78302ea4f148ed0ee0f3ff0b829d1f9c5d

Quantization

  • —Q2-REAP-ds4: compact DS4 profile using IQ2_XXS routed gate/up experts, Q2_K routed down experts, and Q8_0 shared/output/attention projections.

These are DS4/DwarfStar-specific GGUF files for DeepSeek-V4 Flash REAP checkpoints. They are not generic llama.cpp files unless your runtime supports the same DeepSeek-V4 Flash tensor layout and DS4 metadata.

Validation

Validation summaries are uploaded in this repo under:

  • —validation/20260528T160633Z/SUMMARY.md
  • —validation/20260528T160633Z/summary.json

The Spark Q2 GGUF completed the DS4 context sweep through 200000 context on one DGX Spark:

ContextPrefill tok/sDecode tok/sKV bytes
2,048360.2613.6352,184,460
4,096357.0513.7480,373,132
8,192360.2013.56136,750,476
16,384348.3013.31249,505,164
32,768333.7412.59475,014,540
65,536306.4511.79926,033,292
131,072267.6310.291,828,070,796
200,000214.119.252,776,775,308

The corrected 200K API probe used 182,633 prompt tokens and returned the visible marker SPARK-CTX-200000-OMEGA:

Prompt tokensTTFT secondsPrefill tok/sDecode tok/sPassed core
182,633668.53273.1910.61true

Terminal-Bench 2.0 evidence is included in the validation summary: one real gpt2-codegolf trial completed without harness errors after enabling amd64 binfmt on the ARM64 Spark host.

This repo publishes the validated Q2 long-context profile only.

License & citation

License inherited from the base model.

bibtex
@misc{lasby2025reap,
  title  = {REAP the Experts: Why Pruning Prevails for One-Shot MoE Compression},
  author = {Mike Lasby and Ivan Lazarevich and Nish Sinnadurai and Sean Lie and Yani Ioannou and Vithursan Thangarasa},
  year   = {2025}, eprint = {2510.13999}, archivePrefix = {arXiv}
}

Sponsors

Made possible by NVIDIA · TNG Technology · Lambda · Prime Intellect · Hot Aisle.