0xSero/DeepSeek-V4-Flash-180B-GGUF
[!TIP] [Support this work →](https://donate.sybilsolutions.ai) · X · GitHub · REAP paper · Cerebras REAP
DeepSeek-V4-Flash-180B-GGUF
GGUF quantization of 0xSero/DeepSeek-V4-Flash-180B.
At a glance
Which variant should I pick?
This repository contains DS4/DwarfStar GGUF conversions of DeepSeek-V4-Flash-Spark.
The GGUFs point back to the original Spark Hugging Face model:
- Original Spark model: https://huggingface.co/0xSero/DeepSeek-V4-Flash-180B
- Conversion source checkpoint: https://huggingface.co/0xSero/DeepSeek-V4-Flash-180B-codex-K160-REAP
- Runtime/converter repo: https://github.com/antirez/ds4
- Spark deployment repo: https://github.com/0xSero/deepseek-spark
Files
Quantization
Q2-REAP-ds4: compact DS4 profile usingIQ2_XXSrouted gate/up experts,Q2_Krouted down experts, andQ8_0shared/output/attention projections.
These are DS4/DwarfStar-specific GGUF files for DeepSeek-V4 Flash REAP checkpoints. They are not generic llama.cpp files unless your runtime supports the same DeepSeek-V4 Flash tensor layout and DS4 metadata.
Validation
Validation summaries are uploaded in this repo under:
validation/20260528T160633Z/SUMMARY.mdvalidation/20260528T160633Z/summary.json
The Spark Q2 GGUF completed the DS4 context sweep through 200000 context on one DGX Spark:
The corrected 200K API probe used 182,633 prompt tokens and returned the visible marker SPARK-CTX-200000-OMEGA:
Terminal-Bench 2.0 evidence is included in the validation summary: one real gpt2-codegolf trial completed without harness errors after enabling amd64 binfmt on the ARM64 Spark host.
This repo publishes the validated Q2 long-context profile only.
License & citation
License inherited from the base model.
@misc{lasby2025reap,
title = {REAP the Experts: Why Pruning Prevails for One-Shot MoE Compression},
author = {Mike Lasby and Ivan Lazarevich and Nish Sinnadurai and Sean Lie and Yani Ioannou and Vithursan Thangarasa},
year = {2025}, eprint = {2510.13999}, archivePrefix = {arXiv}
}Sponsors
Made possible by NVIDIA · TNG Technology · Lambda · Prime Intellect · Hot Aisle.
