DGX-Spark
SuperQwen3.8-27b-abliterated-NVFP4-DGX-SparkMotif-3-Direct-IQ2-XXS-DGX-SparkDeepSeek-V4.1-Flash-Next-DGX-Spark-512KQwen3.8-Flash-Next-NVFP4-FP8-PBWO-DGX-SparkDeepSeek-V4-Flash-0731-JA-REAP-K216-EXL3-3bpw-DGX-SparkGLM-5.3-Flash-GGUF-DGX-SparkQwen3.8-Flash-Next-NVFP4-Dual-DGX-SparkSuperQwen3.8-Flash-Next-abliterated-FP8-DGX-Spark
dgx-spark-benchmarks
DGX Spark LLM Arena benchmarks
Reproducible LLM inference benchmarks on an NVIDIA DGX Spark (GB10, 128 GB unified memory). The suite defines eleven tests: six closed-loop (llama-benchy) and five open-loop (vllm bench serve). Results cover all eleven: the ten throughput tests under results, and the rate sweep under rateSweep. Raw results remain inspectable, but only complete runs without a failed sanity check count toward rankings and aggregate throughput. Open-loop tests must… See the full description on the dataset page: https://huggingface.co/datasets/Djangodevreng/dgx-spark-benchmarks.dgx-spark-nvfp4-notes
dgx-spark-nvfp4-notes
Assets for the Hugging Face discussion on nvidia/Qwen3.6-35B-A3B-NVFP4: official DGX Spark Marlin recipe dies under concurrent json_schema load on GB10 / SM121.
dgx-spark-benchmarks
DGX Spark LLM Benchmarks
First comprehensive benchmark suite for NVIDIA DGX Spark (GB10 Blackwell).
Hardware
GPU: NVIDIA GB10 Blackwell (1 PFLOP FP4)
Memory: 128GB unified LPDDR5x (273 GB/s)
CPU: 20-core ARM (10x Cortex-X925 + 10x Cortex-A725)
Storage: 4TB NVMe
Framework: Ollama 0.18.3
CUDA: 13.0 | Driver: 580.142
Benchmark Results
Run 1 — General Inference (11 models)
Model
Size
Prompt tok/s
Gen tok/s
Load Time
Llama 3.1 8B
4.9 GB… See the full description on the dataset page: https://huggingface.co/datasets/G3nadh/dgx-spark-benchmarks.glm-5.3-flash-dgx-spark-vllm
GLM-5.3-Flash on 2x NVIDIA DGX Spark (GB10): a working multi-node vLLM config
TL;DR: GLM-5.3-Flash (NVFP4, modelopt quant) running tensor-parallel
across two DGX Spark GB10 nodes over 200GbE RoCEv2 RDMA, with NVFP4
KV-cache, CUDA graphs, and MTP-3 speculative decoding:
~20-23 tok/s single-stream, ~49 tok/s aggregate at 4 concurrent
requests on 64K context. About 2x over a naive TCP/eager launch.
This repo documents the exact configuration, the four vLLM patches it
needs… See the full description on the dataset page: https://huggingface.co/datasets/H-K-B/glm-5.3-flash-dgx-spark-vllm.dgx-spark-kv-cache-benchmark
KV Cache Quantization on NVIDIA DGX Spark GB10
Corrected benchmarks (v3, April 2026) — KV cache quantization behavior on the NVIDIA DGX Spark's GB10 Grace Blackwell unified memory architecture.
Author: Nathan Maine
Date: March 2026, corrected April 2026
Hardware: NVIDIA DGX Spark (GB10, compute 12.1, 128GB unified memory)
Correction Notice: The original v1 benchmarks (March 31) contained methodology errors. Memory was measured via RSS (wrong on unified memory) and some… See the full description on the dataset page: https://huggingface.co/datasets/Nathan-Maine/dgx-spark-kv-cache-benchmark.dgx-spark-moe-benchmarks
Four MoE models on a DGX Spark: speed, tool-calling, and what actually breaks
Full measurement campaign on NVIDIA DGX Spark (GB10, 128 GB unified, ~273 GB/s),
vLLM 0.23.1rc1.dev301+g04c2a8dea, arm64/sm121. Every number here is measured on this
hardware, with the raw evidence included.
The headline: on synthetic tool-calling benchmarks all four models score 91-95 %. In a
real coding agent, three of them score 0-1 out of 14 and one scores 11 out of 14.
If you pick a model from the… See the full description on the dataset page: https://huggingface.co/datasets/pocharlies/dgx-spark-moe-benchmarks.
